More details on Fable 5’s cyber safeguards and our jailbreak framework
Anthropic News ● Covered by 5 sources
Anthropic deployed Claude Fable 5 with new safety classifiers designed to detect and block dangerous cybersecurity uses, while releasing a framework to categorize jailbreak severity. The classifiers sort cybersecurity requests into four categories: prohibited use (ransomware, malware development, data exfiltration), high-risk dual use (penetration testing, exploit development), low-risk dual use (vulnerability identification that other models can already do), and benign use (secure coding, debugging). The framework aims to establish consistent terminology for discussing AI jailbreak risks across government, industry, and academia, with feedback welcomed at cyber-safeguards@anthropic.com and a HackerOne bug bounty program now active.
Why it matters
What is and isn't blocked by our cyber classifiers, and a first draft of our jailbreak severity framework
Also covered by
- The New Stack — Anthropic employees worked “literally around the clock” to keep Fable 5 from disappearing
- Simon Willison — Claude make Fable 5 permanent
- The Batch — Restoration of Claude Fable 5, Gemini's Video Dev Engine, DeepSeek Speeds Up Speculative Decoding
- Anthropic News — Redeploying Claude Fable 5