proomt

Search

Search posts, papers, and topics

All posts

Hugging FaceTuhin Kundu13 min readintermediate

The Agent Said It Was Done. The Database Disagreed.

Summary

ThinkingBox benchmarks AI agents by checking the final database state after tool calls, revealing that many LLM‑driven agents succeed on a single attempt but fail to repeat the correct outcome. Across 507 tasks run 20 times, models differ widely in consistency and cost per successful attempt, with Kimi‑K3 being broad but inconsistent and Claude Opus models being more reliable.

  • Tool calls alone don’t guarantee correct outcomes; the resulting database state must be verified.
  • ThinkingBox runs each task 20 times and reports pass@1, pass@20, and observed 20/20 to measure reliability.
  • High pass@1 scores can hide low consistency—Kimi‑K3 solves 93.9% at least once but only 13.4% consistently.
  • Claude Opus 5.5 improves headline accuracy but adds no consistency over Opus 5.

Teams deploying LLM agents that modify real data need to know not just if they can succeed once, but if they do so reliably and affordably.

7/10

Related reading

  1. Your Agent Aced the Task. Will It Do It Again?

    The post introduces the Consistency Analyzer, a cheap black‑box diagnostic that flags flip‑prone decision steps in LLM agent traces, and shows how feeding the resulting consistency guidelines back into ALTK‑Evolve halves the gap between mean success and all‑run success (Pass⁵) on the AppWorld benchmark without hurting average accuracy.

    Hugging Facehuggingface.co8 minHN21
  2. 2026 in LLMs (so far)

    The post recaps 2026 LLM milestones, noting that Claude Opus 4.5 and GPT‑5.1 made coding agents reliable enough for daily use, sparking AI‑driven side projects and a surge of sandboxing and agent‑security discussions. The author reflects on "AI mania", personal experiments, and the cultural impact on engineers.

    Simon Willisonsimonwillison.net21 minHN5
  3. Constraint Decay: The Fragility of LLM Agents in Backend Code Generation

    A systematic evaluation of LLM agents generating multi‑file backend code shows a sharp drop in correctness when structural constraints (framework conventions, ORM usage, API contracts) are added. Across 100 tasks in 8 Python web frameworks, assertion pass rates fall ~27 points, with data‑layer bugs (bad queries, ORM violations) driving most failures. Mid‑size models cope with minimal frameworks (…

    arXiv cs.SE (Software Engineering)arxiv.org1 minpaperHN287197