TLDRocket
3 October 2026
Microsoft’s ThinkingBox benchmark put a bright, slightly uncomfortable light on what “agent success” really means: not whether a model can describe the right plan, but whether it can reliably reach the correct end state while the database and tool world disagree with the transcript. In a 507-task suite, Microsoft ran 20 repeated executions per task—79,853 attempts failed the executable checks out of 121,680—and among those failures, 77.61% still wrote wrong field values. The headline metric moves from pass-by-appearance toward consistency cost: how often an agent updates the correct records, avoids unintended side effects, and survives the messy reality of gaps between tool-call traces and backend outcomes.
Read the full briefing →