Hugging FaceTuhin Kundu13 min readintermediate
The Agent Said It Was Done. The Database Disagreed.
Summary
ThinkingBox benchmarks AI agents by checking the final database state after tool calls, revealing that many LLM‑driven agents succeed on a single attempt but fail to repeat the correct outcome. Across 507 tasks run 20 times, models differ widely in consistency and cost per successful attempt, with Kimi‑K3 being broad but inconsistent and Claude Opus models being more reliable.
- Tool calls alone don’t guarantee correct outcomes; the resulting database state must be verified.
- ThinkingBox runs each task 20 times and reports pass@1, pass@20, and observed 20/20 to measure reliability.
- High pass@1 scores can hide low consistency—Kimi‑K3 solves 93.9% at least once but only 13.4% consistently.
- Claude Opus 5.5 improves headline accuracy but adds no consistency over Opus 5.
Teams deploying LLM agents that modify real data need to know not just if they can succeed once, but if they do so reliably and affordably.
7/10




