Not all model upgrades are upgrades
TLDR Dev ● Covered by 2 sources
Anthropic's Claude Sonnet 5 has lower per-token pricing but consumes 10x to 12x more tokens on common agent tasks, making it 3.7x more expensive on code upgrades despite the 33% per-token discount. On architecture tasks, Sonnet 5 produced lower quality output (78% vs 90% on idiomatic evaluation) while consuming vastly more tokens, though it excelled at precise instruction-following on code upgrade tasks. Neither model overcomes undocumented information gaps, so improving agent grounding content delivers more value than upgrading models.
Why it matters
Not all upgrades to AI models result in improvements. A comparison of two models, Claude Sonnet 4.6 and Claude Sonnet 5, showed that the newer model can incur higher costs and produce inferior output in certain tasks despite cheaper token pricing.