Testing Mythos and Fable, Moving Beyond SWE-bench, Nvidia's Open Contender
The Batch ● Covered by 17 sources
Anthropic and the US government both showed they can flip off access to frontier AI models overnight, sparking new benchmark chaos and a scramble for alternatives. Meanwhile SWE-bench is losing relevance as harder new coding tests emerge.
Based on reporting by The Batch — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Andrew Ng's latest letter isn't really about a product launch. It's about power, and who gets to pull the plug. Over two weeks, Anthropic released Claude Fable 5, a guardrailed version of its Mythos model, then quietly degraded performance for users it detected were doing LLM research, then walked that back under backlash but kept refusing to let the model help AI researchers outright. Days later, the U.S. Commerce Department stepped in, requiring export licenses for any foreign national to use Claude Mythos 5 or Claude Fable 5 — even Anthropic's own overseas employees. Anthropic responded by shutting off Claude Fable 5 access worldwide.
Ng's real complaint is about competitive terms baked into Fable 5's usage rules: developers can't use it to build rival LLM technology, a restriction he compares to Microsoft banning competitive software built on its tools. He also flags a new wrinkle — a mandatory 30-day data retention policy for anyone using Fable 5 — as evidence that building on any single proprietary provider now looks shakier than it used to.
That instability shows up concretely in how the model gets tested. Because Fable 5's classifiers screen every prompt before it reaches the model, flagged questions either get silently rerouted to the weaker Claude Opus 4.8 or refused outright, depending on whether you're using Anthropic's own app or the API. Evaluators like Artificial Analysis, Vals AI, and the team behind Agents' Last Exam all had to decide whether to score Fable 5
My take — AI-written commentary, not fact-checked reporting
Export controls dressed up as safety theater are exactly the kind of move that backfires slowly and then all at once — nations don't forget being cut off, they just quietly build workarounds, the way China did with chips once it got squeezed. Anthropic's competitive-use clause is the tell here: if the concern were really bioweapons and hacking, nobody would need a rule banning rival LLM builders from your API. The open-source crowd should be thrilled, because every one of these moves is a recruitment poster for training your own model instead of renting someone else's.
Read more about this at: The Batch