Anthropic's Safety Superpower
stratechery.com ● Covered by 3 sources
Anthropic released Fable, a powerful new AI model, then the US government forced them to pull it days later over a jailbreak scare. Turns out safety concerns, data grabs, and Anthropic's rivalry with rivals might all be tangled together.
Based on reporting by stratechery.com — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Two months ago Anthropic said its Mythos model was too dangerous to release. Then it shipped a guardrailed version called Fable anyway — and the author, after using it, says it made GPT-5.5 and Opus 4.8 feel small by comparison, the kind of leap he'd only felt before with GPT-4 and Grok 4. Shortly after launch, someone found a jailbreak. The US government responded with an export control order forcing Anthropic to disable Fable and Mythos for every foreign national, including its own employees, while Anthropic scrambled to Washington insisting this was a misunderstanding over what it says are minor, already-known vulnerabilities.
The jailbreak fight is almost beside the point. What's more interesting is why Anthropic released a model it called dangerous in the first place, and the answer is money. Compute has captured most of AI's economic value so far — Nvidia, TSMC, SK hynix, Samsung, Micron — while OpenAI and Anthropic burn tens of billions building frontier models that open-source competitors, often from China, quickly commoditize. The way out of that trap is to stop being a commodity input and start owning the user relationship directly, which puts the labs on a collision path with the software industry itself. Satya Nadella's recent essay warning against a future where a handful of models
My take — AI-written commentary, not fact-checked reporting
,
Read more about this at: stratechery.com