TLDRocket
Sign in

Claude Sonnet 5.5 vs. Opus 5.5: 42% cheaper and perfect on every run

The New Stack Jessica Wachtel ● Covered by 15 sources

Anthropic’s Sonnet 5.5 just beat Opus 5.5 in a head-to-head test. It was perfect on every run and ended up about 42% cheaper overall.

Based on reporting by The New Stack, Jessica Wachtel — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Anthropic shipped Sonnet 5.5 just six days after Opus 5.5, and the pitch is familiar: lower price, strong benchmark numbers, faster output. On Terminal-Bench 4.0, the company says Sonnet 5.5 scores 70.6%, ahead of Opus 5.5’s 66.4% at xhigh effort. It also says Sonnet 5.5 produces output more than 30% faster than Sonnet 5 and uses fewer tokens per task.

But the real story here is that token pricing and actual bills are not the same thing. Sonnet 5.5 is priced at $2 per million input tokens and $10 per million output tokens, exactly half of Opus 5.5’s $4 and $20. In practice, though, a model that burns through more tokens can erase that discount fast.

That is what independent testing from Artificial Analysis found at max effort, where Sonnet 5.5 cost $7.67 per task versus $5.98 for Opus 5.5. The New Stack’s own testing went the other way. Across three coding-style tests run five times each through the Anthropic API, with adaptive thinking and maximum effort enabled, Sonnet 5.5 was perfect on all 15 runs and Opus 5.5 was perfect on 13.

The tests covered an agentic bug fix task, a dependency resolver spec, and an asyncio concurrency cleanup. Sonnet 5.5 won the resolver spec and the concurrency bugs, while Opus 5.5 won the agentic bug fix. That first test is also where Sonnet’s appetite for thinking caused trouble: on the initial attempt it hit a 32,000-token step limit in four of five runs, and those runs had to be redone after the limit was raised to 128,000.

Across all 15 runs, Sonnet 5.5 cost $12.69 and finished in 2:11:36, compared with Opus 5.5’s $22.07 and 2:28:39. Counting the reruns, Sonnet still came out about 36% cheaper overall, not half off, because it often used more tokens to get there. On the hardest coding tasks, that still looks like a useful default. Just not a magical one.

My take — AI-written commentary, not fact-checked reporting

This is the usual AI pricing trap: cheap per token looks neat until the model starts thinking like it’s billing by the hour. Sonnet 5.5 sounds like the better default for hard coding, but the agent loop crowd will keep paying for Opus when speed matters. The industry loves simple price cuts; the bill, annoyingly, still has opinions.

Read more about this at: The New Stack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.