The Batch
● 5 sources
Claude Fable 5 and Mythos 5 models were restored by Anthropic on July 1 after a three-week suspension imposed by the U.S. government over national security concerns related to cybersecurity capabilities. Anthropic implemented additional guardrails that route certain cybersecurity queries to the less capable Claude Opus 4.8, and the models are now available through the Claude API and other Anthropic platforms. The incident marks the first time a government intervention led to suspension of general access to an AI model, likely to influence how future models are reviewed and released by other AI companies.
Anthropic News
● 5 sources
Anthropic deployed Claude Fable 5 with new safety classifiers designed to detect and block dangerous cybersecurity uses, while releasing a framework to categorize jailbreak severity. The classifiers sort cybersecurity requests into four categories: prohibited use (ransomware, malware development, data exfiltration), high-risk dual use (penetration testing, exploit development), low-risk dual use (vulnerability identification that other models can already do), and benign use (secure coding, debugging). The framework aims to establish consistent terminology for discussing AI jailbreak risks across government, industry, and academia, with feedback welcomed at cyber-safeguards@anthropic.com and a HackerOne bug bounty program now active.
Anthropic News
● 5 sources
Anthropic restored access to Claude Fable 5 and Mythos 5 after the US government lifted export controls that had been imposed on June 12 following a jailbreak vulnerability discovered by Amazon researchers. The new safety classifier blocks the reported bypass technique in over 99% of cases, though it increases false positives during routine coding tasks. Fable 5 becomes available globally starting July 1, with Anthropic, Amazon, Microsoft, Google and others now developing a shared industry framework for assessing AI jailbreak severity to standardize future responses.