TLDRocket
Sign in

You picked Claude Sonnet 5.5 — but Anthropic may send your request to Sonnet 5 in “higher-risk” situations

The New Stack Amanda Caswell ● Covered by 13 sources

Anthropic’s Claude Sonnet 5.5 can quietly hand risky cyber requests to Sonnet 5. That means the “new” model isn’t always the model answering.

Based on reporting by The New Stack, Amanda Caswell — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Anthropic has turned Claude Sonnet 5.5 into a test case for a new kind of safety routing: the cheaper Sonnet tier now gets the same style of cyber safeguards and fallbacks the company has been using on its strongest models. The twist is that Sonnet 5.5 isn’t Anthropic’s most capable model overall, but it is strong enough in cybersecurity to trigger those controls anyway.

The company says the model does not push its frontier capabilities forward, yet it puts up serious numbers on offensive security work. With cyber safeguards disabled, Sonnet 5.5 got full arbitrary code execution in 178 of 410 ExploitBench runs, completed 46.1% of CyScenarioBench, and reached 50 control-flow hijacks on a benchmark built from Google’s OSS-Fuzz corpus. Anthropic still rates it below Opus 5.5 and Mythos 5.1 for cybersecurity, but the jump from Sonnet 5 was big enough to move it under the same policy umbrella.

That policy is not just a simple block button. Anthropic says enforcement starts with a probe over the model’s activations, then a lightweight classifier on Sonnet 5.5, and then a separate trained LLM classifier that helps decide whether a conversation gets stopped. In some higher-risk cybersecurity cases, the request can be sent to Sonnet 5 instead. Anthropic says that can include penetration testing, exploit generation, and binary vulnerability scanning. Its own docs also say routine software development is unaffected.

But the catch is obvious: this is not a clean model swap. Anthropic’s apps automatically route blocked cyber requests to Sonnet 5, while API users have to opt in, and blocked requests in the API just stop if fallback is off. The checks also look at memory, connector content, web search results, and files, so code that never came from the user can still trip the system. And because fallback can move a request to an older model, Anthropic’s own prompt-injection tests found that 25% of requests sent to Sonnet 5.5 were rerouted to Sonnet 5, with 12.01% of those rerouted requests compromised.

Anthropic is still tuning the classifiers to cut false positives, and it plans to widen access for verified defenders through its Cyber Verification Program. For now, though, Sonnet 5.5 is a reminder that “safer” AI often means “more routing,” not fewer surprises.

My take — AI-written commentary, not fact-checked reporting

Anthropic keeps proving the same awkward point: once a model gets useful enough for serious security work, the neat idea of a single answer model falls apart. Safety here is really a traffic system, and traffic systems add friction, false stops, and the occasional detour through an older engine. That’s not a bug so much as the price of shipping powerful models without pretending they’re harmless toys.

Read more about this at: The New Stack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.