Anthropic’s Smaller Model Sonnet 5.5 Overtakes OpenAI’s Flagship GPT-6
Trending Topics Jakob Steinschaden ● Covered by 5 sources
Anthropic’s new Sonnet 5.5 beats OpenAI’s top GPT-6 models on an independent A.I. chart. The catch: it does it with a huge token bill.
Based on reporting by Trending Topics, Jakob Steinschaden — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Anthropic has put out Claude Sonnet 5.5, and the smaller model has jumped straight to No. 2 on Artificial Analysis’s Intelligence Index. It scored 56 points, just two behind Anthropic’s own Opus 5.5. OpenAI’s GPT-6 Astra sits at 53, with GPT-6 Sol at 48.
That makes Anthropic the only company with the top two spots. Sonnet 5 is not a small upgrade over Sonnet 5 either; the new model gains 18 points over its predecessor’s 38. On a benchmark that mixes coding, knowledge questions and agent-style tasks, that is a serious climb.
The model looks especially strong when it has to do work on its own. On Terminal-Bench 4.0, Sonnet 5.5 reaches 64 percent, ahead of both Opus 5.5 and GPT-6 Astra at 60 percent. On Terminal-Bench-Science, it lands at 53 percent, behind only those two. For everyday office work, it is almost neck and neck with Opus: 1,844 versus 1,846 on GDPval-AA, 1,811 versus 1,822 on AA-Briefcase, and 71 versus 70 percent on AutomationBench-AA.
But there is a bill attached to that performance. In its highest reasoning mode, Sonnet 5.5 uses about 193,000 output tokens per task, the most Artificial Analysis has ever measured. That is roughly 60 percent more than Opus 5.5 and about seven times GPT-6 Astra. Anthropic kept the sticker price at $2 per million input tokens and $10 per million output tokens, but a single Intelligence Index task still costs about $7.60 — more than Sonnet 5 and far above GPT-6 Sol’s $1.06.
Anthropic says the public model fixed a bug that affected structured outputs in the pre-release version Artificial Analysis tested, so results should not move much. Sonnet 5.5 is already in the Claude apps, the Anthropic API, and on AWS, Google Cloud and Microsoft Azure. And while GPT-6 Luna is much cheaper at 10 cents per million input tokens, Anthropic clearly wants the market to focus on the top of the chart, not the bargain bin.
My take — AI-written commentary, not fact-checked reporting
This is the old AI story in sharper clothes: the best model is rarely the one you can afford to run all day. Anthropic can claim the crown, but the token appetite is doing some embarrassing work in the background. The industry keeps treating efficiency like a footnote until the bill arrives.
Read more about this at: Trending Topics