TLDRocket
Sign in

Claude Opus 5.5 vs. Opus 5 on reasoning tasks: Cheaper, faster, but not better

The New Stack Jessica Wachtel ● Covered by 6 sources

Claude Opus 5.5 is cheaper and faster than Opus 5, but it didn’t beat it on the reasoning tests here. The surprise: the savings came from fewer tokens, not better answers.

Based on reporting by The New Stack, Jessica Wachtel — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Anthropic says Claude Opus 5.5 is a step up from Opus 5: cheaper, faster, and good enough to match its stronger sibling’s level. On paper, the company’s pitch is tidy. In practice, a set of reasoning-only tests from The New Stack shows something messier: the new model is mostly the same brain with a better bill.

The pricing change is real. Through the API, Opus 5.5 costs $4 for a million input tokens and $20 for a million output tokens, down from $5 and $25 for Opus 5. Anthropic says that should translate to 40% lower costs on typical workloads, and that the model writes output more than 30% faster. The tests here didn’t use the usual developer-style simulations. They focused on three hard reasoning tasks: a logic grid, a constrained ordering puzzle, and a stone game.

On the logic grid, both models got all 28 cells right. Opus 5.5 finished in 65 seconds and spent $0.16, while Opus 5 took 108 seconds and cost $0.27. That’s a solid win for the newer model, and the cheapest kind too: same answer, less waiting, less money.

The ordering puzzle was a different story. Neither model produced an answer at the 48,000-token limit, and even after the limit was raised to 128,000 tokens, both still failed. Opus 5.5 hit a refusal stop after 18 minutes and 56 seconds; Opus 5 ran for 25 minutes and 24 seconds before stopping with no answer. The prompt wasn’t sensitive, so the refusal looks like a safety filter mistake. Either way, the result was the same: a lot of billed thinking, no useful output.

The stone game put the models back on level ground for correctness, but not for efficiency. Both solved it. Opus 5.5 did so in 215 seconds and $0.58, while Opus 5 took 624 seconds and $1.88. Across all the calls in this test set, Opus 5.5 wrote at 103.4 tokens per second versus 93.1 for Opus 5, about 11% faster overall. That’s still well short of Anthropic’s 30% claim. The bigger gain came from using fewer tokens, which is why the bill dropped so much.

So the practical takeaway is pretty blunt. If someone is already paying for Opus 5, Opus 5.5 looks like the better buy. But this wasn’t a story about smarter reasoning; it was a story about lower spend, shorter waits, and the occasional expensive nothing.

My take — AI-written commentary, not fact-checked reporting

This is the usual AI trade: fewer tokens, nicer invoice, same old probability theater. I’m pro-closed models when they’re better, not when the pitch deck says “faster” and the hardest tasks still end in silence. If a model needs a safety filter apology for counting jobs, the market should treat that as a bug, not a feature.

Read more about this at: The New Stack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.