The Agent Said It Was Done. The Database Disagreed.
Hugging Face
Microsoft’s ThinkingBox grades AI agents by what they change in the database, not what they say. That matters because a polished answer can still leave the real record wrong.
Based on reporting by Hugging Face — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Microsoft’s ThinkingBox is trying to answer a blunt question: did the agent actually finish the job, or just talk like it did? The benchmark, now available through Hugging Face, judges agents on backend state and side effects instead of on the neatness of their responses. It also runs each task 20 times, because one lucky pass doesn’t prove much.
The paper’s retail example makes the point without any theory-wrangling. A customer has a $745 kitchen appliance stuck in a courier exception at a Nashville distribution center, fifteen days past the estimated delivery date. The agent does the polite, thorough thing: looks up the order, checks tracking, opens a ticket, reads the policy, and even gets the policy right. But it closes the ticket as resolved when the carrier exception is still open, and the customer still hasn’t gotten a real answer to the question she asked.
That kind of failure is exactly what ThinkingBox is built to catch. Across 507 stateful business workflows, each run 20 times against different models, the benchmark checks whether the terminal database state matches the required end state. In the common-set ablation, 79,853 of 121,680 valid trials failed the executable checks. More than two-thirds of those failures still ended cleanly, called a state-changing tool, and reported no final tool error. Yet the database still found wrong field values in 77.61% of them, extra effects in 43.30%, and missing required effects in 25.36%.
The headline model numbers are a reminder that single-shot scores can flatter the wrong thing. Claude Opus 5.5 leads overall on pass@1 at 67.16%, just ahead of Claude Opus 5 at 66.50%. Kimi-K3 is the strongest open-weight model on the single-attempt metric at 57.37% overall, and it reaches 93.89% of tasks at least once. But consistency tells a different story. Kimi-K3 passes only 68 tasks on all 20 runs, while Claude Opus 5 completes 241 tasks on every run. Claude Opus 5.5 improves the headline score again, but it still matches Claude Opus 5 on the all-20s count: 241.
Cost doesn’t rescue the easy story either. GPT-5.6 Sol is cheapest per successful task attempt at $0.127, but GPT-5.4 and Claude Opus 5.5 form the better bargain when consistency is the metric, with GPT-5.4 at $6.80 per dependable task and Claude Opus 5.5 at $7.80. The benchmark’s own breakdown says most failures are tool-handling problems, not pure reasoning mistakes. That is a pretty useful reminder: agents usually don’t fail by being confused poets. They fail by not making the right thing happen.
My take — AI-written commentary, not fact-checked reporting
This is the right way to test agents, because a database does not care how eloquent the assistant sounded. The industry has spent a long time grading the costume and calling it performance. ThinkingBox pulls the curtain and finds the usual thing: lots of confident nonsense with a transaction log attached.
Read more about this at: Hugging Face