TLDRocket
Sign in

Claude did best on a new benchmark for ‘agents that build agents’. It still passed fewer than a quarter of the tests.

The New Stack Paul Sawers Covered by 2 sources

Claude topped Sierra’s new test for AI that builds other AI, but it still cleared only 23.9%. That gap says agent-building is still way harder than vendor demos make it look.

Based on reporting by The New Stack, Paul Sawers — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Sierra has opened up a nastier version of its earlier agent benchmark: instead of checking whether an AI can serve customers, Hyper-τ-bench asks whether an AI can build the customer-service agent in the first place. That sounds neat until the scores arrive. The best result in Sierra’s first public run was Claude Opus 5 in Claude Code, at 23.9%. GPT-5.6 Sol in Codex was close behind at 22%. None of the six autonomous setups made it past 25%.

The test is deliberately messy. Sierra hands the developer agent a simulated business with documents, transcripts, an API and a codebase, then asks it to create a customer support agent under model and cost limits. The finished system is judged on unseen airline, retail, telecom and banking cases — things like canceling a flight or disputing a fee — and it only passes if it gives the right answer and makes the right change in the underlying systems.

The low scores are not the whole story. Sierra also published a human-plus-AI reference line at 82.2%, but the company is careful not to oversell it as ordinary human performance. Those reference builds were made by a benchmark author working with a frontier model and access to the ground-truth requirements, which the autonomous agents did not get. In other words, this is an oracle, not a fair fight.

What matters more is where the autonomous builders fell apart. They often quit researching too early, asked too few questions, and settled on the first architecture that worked instead of trying alternatives. Banking was especially brutal: Claude Opus 5 reached 5.9% there, while GPT-5.6 Sol got 9%. That domain carries 35 of the benchmark’s 53 tasks and depends on 2,969 policy facts, with some tasks touching as many as 580 of them.

The pattern gets even more human in the failure mode. Sierra says the agents opened fewer than 80 of about 1,700 available banking files and leaned on search instead of digging through the material. Across the logged runs, question-asking accounted for just 0.3% of tool calls. And when Sierra nudged one telecom build with a single sentence about a different architecture, its score jumped from 31% to 67%. The machines aren’t helpless; they’re just still very easy to steer off a cliff by not giving them enough context.

My take — AI-written commentary, not fact-checked reporting

This is the kind of benchmark that quietly ruins the marketing deck, which is exactly why it matters. Agent builders keep acting like one more wrapper will fix judgment, but the test says the hard part is still the boring, stubborn work of gathering facts, asking questions, and not shipping the first thing that runs. That’s less glamorous than “autonomous developer,” but a lot closer to reality.

Read more about this at: The New Stack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.