TLDRocket
Sign in

Claude

91 summarised stories about Claude, each linking back to the original source. Browse all topics →

+ Follow this topic

Saturday, 25 July 2026

Claude Opus 5: The System Card

Zvi (Don't Worry About the Vase) 1 month ago 2 17 sources

Anthropic released Claude Opus 5, a model positioned between Opus 4.8 and their larger Mythos 5, claiming comparable performance to Fable 5 at half the price and faster speed. Key concrete improvements include reducing prompt injection attack success rates from 5.5% to 2.0% on the IPI benchmark, cutting safety classifier false positives from 42% to 5% in FrontierBench, and achieving 69% multi-turn appropriate response rates for self-harm queries versus 58% for Fable. The model now permits vulnerability analysis in source code while maintaining blocks on binary analysis, and shows improved agentic safety and alignment compared to previous versions, though it lacks the full multi-step exploit capability of Mythos 5.

Quoting Boris Cherny

Simon Willison's Weblog 1 month ago 19 17 sources

Anthropic engineer Boris Cherny stated that Claude Opus 5 is their most resistant model to prompt injection attacks. The claim is documented in the model's system card on page 73, with results from prompt injection evals and red teaming across their safety testing. This suggests Opus 5 offers improved robustness against a common method of manipulating AI model behavior.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.