TLDRocket
Sign in

SOP-Bench: A new benchmark for evaluating AI agents on real business procedures

Amazon Science

Amazon launched SOP-Bench, a benchmark for AI agents to follow real business procedures. It tests them on messy, tool-heavy tasks, not just neat text answers.

Based on reporting by Amazon Science — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Amazon Science has put out SOP-Bench, a benchmark built to see whether AI agents can actually carry out real standard operating procedures instead of just sounding smart about them. The target is the kind of work organizations run on every day: patient intake, dangerous-goods checks, customer service, content moderation, financial compliance, warehouse inspection. The twist is that the procedures are left as humans wrote them, with the gaps, ambiguities, and assumptions intact.

That matters because SOPs are exactly where polished demos tend to crack. A procedure may say “verify insurance” twice, but a seasoned employee knows those are two different checks. A model does not get that background for free. It has to infer the right meaning, keep track of earlier steps, and choose between tools that may look similar but do very different things.

SOP-Bench is meant to close a gap in existing agent testing. Amazon says it is the first benchmark to combine genuine enterprise procedures, working tools, and ground-truth answers so an agent is scored on completing the task, not on producing text that flatters an automated judge. The release includes more than 2,000 tasks across 12 business areas, along with tool interfaces and correct outcomes. It also ships as a framework, so teams can plug in their own agents and even add their own procedures.

The benchmark was presented at the 2026 Conference on Knowledge Discovery and Data Mining, and Amazon says it was built with domain experts and Anthropic Claude 3.5 Sonnet v2 doing the mechanical conversion work. The experts wrote the original procedures, checked the logic, verified the data, and reviewed the generated code. Amazon says no proprietary or sensitive data was used.

The early results are not kind to the usual “just upgrade the model” instinct. On one reasoning-style setup, Claude 4.5 scored lower than Claude 4. Adding tools can also backfire: in one video-annotation procedure, performance nearly halved when the required six tools were buried inside a larger set of 26. And there was no universal winner. Some tasks were easy, like triaging emails by intent, where agents were right about nine times out of ten. Others, like annotating objects in a driving video, were right only about one time in four.

My take — AI-written commentary, not fact-checked reporting

This is the kind of benchmark the field badly needs, because agents keep being sold like they’re general workers when they’re really flaky interns with a toolbelt. The awkward part is that the messier the real procedure gets, the less useful the leaderboard fantasy becomes. Benchmarks that force models to survive actual operations are the only ones worth the oxygen.

Read more about this at: Amazon Science

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.