Claude Opus 5.5: Anthropic Launches New Top Model Despite Calling for AI Slowdown
Trending Topics Jakob Steinschaden ● Covered by 2 sources
Anthropic just shipped Claude Opus 5.5, and it now tops a key benchmark. Odd timing: the launch lands days after its CEO called for slower AI progress.
Based on reporting by Trending Topics, Jakob Steinschaden — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Anthropic has released Claude Opus 5.5, the first model in its new Claude 5.5 family, and it arrives with a contradiction baked in. The company’s own CEO, Dario Amodei, had just argued in a widely discussed essay that the industry should slow down. Yet here is a new flagship model, cheaper than before and, by one independent benchmark, better than anything else tested so far.
Artificial Analysis puts Opus 5.5 at the top of its Intelligence Index with a score of 58 at maximum effort. The firm says that is the highest result it has measured, and it says the model leads six of the ten evaluations behind the index. That includes Humanity’s Last Exam, where it scored 61.4 percent, and SciCode, where it reached 66.9 percent. On Terminal-Bench 4.0, Artificial Analysis measured 59.6 percent, which leaves Anthropic level with OpenAI’s GPT-6 Astra in that test and 11 points ahead of Opus 5.
The model also looks strong on work that resembles real office output rather than neat benchmark tricks. On Artificial Analysis’s private AA-Briefcase test, it reaches an Elo rating of 1,822, 143 points above Claude Fable 5.1. The testers say this is the first time an Anthropic model has beaten OpenAI’s GPT-5.6 Sol on presentation quality. On GDPval-AA v2.1, which spans 44 occupations, Opus 5.5 scores 1,846 Elo.
Price cuts are the other half of the story. Anthropic has lowered input tokens to $4 per million and output tokens to $20 per million, while cache reads drop to $0.20 per million tokens. The company says typical workloads should end up 40 percent cheaper. But Artificial Analysis adds a catch: at maximum effort, Opus 5.5 uses about 119,000 output tokens per task, far more than Opus 5 and well above GPT-6 Astra. That leaves the new model about even on cost per task rather than obviously cheaper.
Not everything moves forward. Artificial Analysis says Opus 5.5 trails rivals on CritPt, AA-LCR and GDP.pdf. Anthropic’s own table also shows GPT-6 Astra ahead on Terminal-Bench-Science and narrowly ahead on AutomationBench. And then there are the new constraints: some cybersecurity and biology work gets routed to older models, full access for sensitive tasks requires verification, and thinking mode can’t be switched off. Anthropic says the model is now available in the Claude apps, on the Claude Platform, and through AWS, Google Cloud and Microsoft Azure.
My take — AI-written commentary, not fact-checked reporting
This is the classic frontier-model move: call for restraint, then ship the thing anyway and hope the press release does the moral heavy lifting. Anthropic clearly wants the credit for caution and the revenue from shipping, which is a very Silicon Valley compromise. The industry keeps treating “pacing” like a public mood and not an actual brake.
Read more about this at: Trending Topics