TLDRocket
Sign in

Anthropic Releases Claude Sonnet 5.5: 70.6% on Terminal-Bench 4.0 at the Same $2/$10 Price

MarkTechPost Michal Sutter ● Covered by 12 sources

Anthropic launched Claude Sonnet 5.5, a faster, cheaper sibling to Opus 5.5. It keeps the same $2/$10 price, but says it burns far fewer tokens per task.

Based on reporting by MarkTechPost, Michal Sutter — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Anthropic has launched Claude Sonnet 5.5, the second model in its Claude 5.5 line after Opus 5.5. The pitch is straightforward: give people something quicker and cheaper for everyday work, while leaving the heavier lifting to Opus when needed.

Sonnet 5.5 is aimed at well-scoped tasks, bug fixes, and producing clean documents, slides, and spreadsheets. It’s live now on the Claude Platform as claude-sonnet-5-5, and it’s also available through AWS, Google Cloud, and Microsoft Azure. But it’s still a closed-weights model, so self-hosting is off the table.

Compared with Sonnet 5, Anthropic says output generation is more than 30% faster, cost per task is up to 30% lower, and the writing is clearer. It also says this is the first Sonnet model to beat Pokémon Red using only screenshots, which is a wonderfully odd way to measure progress and exactly the kind of benchmark AI teams keep inventing when the usual ones start to feel stale.

The model comes with a 1 million-token context window, 128K max output, and a June 2026 reliable knowledge cutoff. Adaptive thinking is on by default, with five effort levels: low, medium, high, xhigh, and max. Anthropic says Claude Code and the Claude apps default to Medium, while the Claude Platform defaults to High.

On benchmarks, the headline number is Terminal-Bench 4.0, where Sonnet 5.5 scored 70.6%, far ahead of Sonnet 5’s 10.3% and above Opus 5.5’s 66.4% at Xhigh. It also posted 55.5% on CursorBench 4.0, 52.1% on FrontierCode 1.1 at Xhigh, 80.1% on OSWorld 2.1, and 64.5% on Humanity’s Last Exam with tools. Anthropic says the Max setting can score lower than Xhigh because the model more often runs multi-agent code review, which can trigger timeouts or out-of-scope edits.

The sticker price hasn’t moved: $2 per million input tokens and $10 per million output tokens, with cache reads at $0.20 and cache writes at $2.50 per million. The savings, Anthropic says, come from efficiency, not a price cut. Customer examples back that up: Balyasny Asset Management saw about 121K tokens per answer versus 497K on Sonnet 5, Base44 reported 3.6 iterations per app build instead of 7.7 with Opus 5, and Zendesk said tickets were processed 20% faster.

My take — AI-written commentary, not fact-checked reporting

Closed models keep getting sold as convenience, and yes, they are convenient. But the real story here is how much of the cost win comes from better token discipline, not cheaper pricing — the cloud gets to keep the toll booth, and everyone claps anyway.

Read more about this at: MarkTechPost

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.