TLDRocket
Sign in

Quoting Boris Cherny

Simon Willison Simon Willison Covered by 13 sources

Anthropic engineer Boris Cherny stated that Claude Opus 5 is their most resistant model to prompt injection attacks. The claim is documented in the model's system card on page 73, with results from prompt injection evals and red teaming across their safety testing. This suggests Opus 5 offers improved robustness against a common method of manipulating AI model behavior.

Why it matters

More than any of these eval scores, what is most exciting to me is something else: Opus 5 is our least prompt injectable model yet. It is a bit buried in the system card, but across PI evals and red teaming, Opus 5 is very hard to prompt inject successfully. — Boris Cherny, here's that System Card section, page 73 Tags: prompt-injection, anthropic, claude, generative-ai, ai, llms, boris-cherny

Also covered by

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.