Introducing Claude Sonnet 5
Anthropic ● Covered by 3 sources
Anthropic just launched Claude Sonnet 5, its most agentic Sonnet model yet, closing in on Opus-level performance for less money. It handles multi-step coding and tool-use tasks that used to need pricier models.
Based on reporting by Anthropic — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Anthropic has shipped Claude Sonnet 5, and the pitch is straightforward: this is the Sonnet line finally catching up to what Opus-class models have been doing, but at a fraction of the price. For a while now, the biggest agentic gains have come from the bigger, pricier Opus models. Sonnet 5 is Anthropic's attempt to close that gap, and by the company's own account, its performance lands close to Opus 4.8 on reasoning, tool use, coding, and knowledge work, while costing noticeably less to run.
The model launches with introductory pricing of $2 per million input tokens and $10 per million output tokens, holding through August 31, 2026, before rising to $3 and $15. It's live everywhere at once: default for Free and Pro users, available to Max, Team, and Enterprise plans, and accessible through Claude Code and the Claude Platform via the API. Anthropic also bumped up rate limits across Chat, Cowork, Claude Code, and the Platform to handle the extra token usage that comes with running the model at higher effort levels.
What early access partners describe is less about raw benchmark numbers and more about follow-through. Testers quoted in Anthropic's announcement talk about the model finishing jobs that previously stalled halfway, checking its own work without being told to, and tracing bugs to root causes instead of patching symptoms. One partner described handing it a two-part task, updating Salesforce tiers and sending a launch email, and having it complete both ends without intervention. Another had it investigate a bug and, unprompted, write a reproducing test, apply a fix, then confirm the bug returned without the change. That kind of self-checking behavior is the throughline across the feedback Anthropic chose to publish.
On safety, Anthropic says Sonnet 5 shows an overall lower rate of undesirable behavior than its predecessor, Sonnet 4.6, including better resistance to prompt injection and hijack attempts, and lower rates of hallucination and sycophancy. It still scored somewhat worse on Anthropic's automated behavioral audit than the more capable Opus 4.8 and Claude Mythos Preview. Notably, the company says it did not deliberately train Sonnet 5 on cybersecurity tasks, and testing against Firefox vulnerabilities found it couldn't produce a working exploit at all, though it showed a slightly higher rate of partial success than Sonnet 4.6, something Anthropic chalks up to general intelligence gains rather than targeted training. Because of that modest uptick, Sonnet 5 ships with the same cyber safeguards used in Opus 4.7 and 4.8, though Anthropic still points enterprise customers needing looser guardrails for cybersecurity work toward Opus 4.8 instead.
My take — AI-written commentary, not fact-checked reporting
The real story here isn't the raw scores, it's the pricing move: squeezing near-Opus agentic performance into the cheaper Sonnet tier is exactly the kind of pressure that keeps the whole field honest, forcing every other lab to justify why their mid-tier models still cost more for less. Anthropic deserves some credit for being upfront that Sonnet 5 lags behind Opus and Mythos Preview on safety audits and cyber capability, rather than burying the caveat. That kind of transparency should be the baseline, not a selling point.
Read more about this at: Anthropic
Related stories
Meet the New Claude Opus 5: Frontier-Class Agentic Coding and Computer Use at Unchanged Opus Pricing
MarkTechPost · 1 month ago ·
34
Introducing Claude Opus 5
Anthropic ·
29