Anthropic released Claude Opus 5, replacing Claude Opus 4.8 at the same price of $5 per million input tokens and $25 per million output tokens. The model enables thinking by default (whereas Opus 4.8 required manual activation), supports up to 1 million token context with 128k output on the standard API, and requires developers to remove verification prompts from existing code since the model now self-verifies. Developers must update API calls to handle the new thinking behavior or explicitly disable it, and they can reduce prompt cache minimums from 1,024 to 512 tokens.
Anthropic released Opus 5, an updated language model focused on cost efficiency rather than major capability gains. The model performs at roughly the same level as Anthropic's Fable on coding benchmarks while costing approximately half as much, and shows only iterative improvements over Opus 4.8. Opus 5 deliberately lacks some cybersecurity training and drops certain safety policies, positioning it as a practical middle ground for developers prioritizing affordability over cutting-edge performance.
Anthropic released Opus 5, a model smaller than its Fable 5 but cheaper and less restrictive while outperforming Fable 5 on several benchmarks. The model launched two months after Opus 4.8, and Anthropic reports its safety classifiers will engage 85% less often for Opus 5 than for Fable 5. Users now have access to a model with fewer restrictions and a new Automatic Fallbacks feature that routes blocked requests to a weaker model instead of returning an error.
The article guides users on which AI systems to use for different tasks, explaining that agentic systems now let AI handle multi-hour work autonomously with computer access, going beyond simple chatbot conversations. Claude Opus and ChatGPT's GPT-5.6 Sol at "High" thinking levels are recommended for high-stakes work, while ChatGPT Work and Claude Cowork (with company-provided computers) or Codex and Claude Code (with personal computer access) enable the most powerful applications. The key difference is that desktop versions grant the AI access to your actual computer, enabling complex multi-file projects like the author's book fact-checking task that took 30 minutes and verified 195 references without errors, fundamentally changing AI from a chatbot interface to a delegated work team.
Anthropic expanded Claude's voice mode to run on its more capable Opus and Sonnet models, not just Haiku, while adding support for 12 languages and integration with tools like Gmail and Slack. Voice mode now defaults to the fastest version of whichever model users select, allowing mid-conversation switching between Haiku, Sonnet, and Opus for deeper reasoning on complex problems. Users can now have longer, more sophisticated voice conversations for business decision-making and can trigger connected tools through voice commands after requesting permission.
Researchers identified a "no-recovery bottleneck" in large language models attempting long-horizon reasoning tasks, where errors on difficult steps become irreversible despite task decomposition. They developed Lookahead-Enhanced Atomic Decomposition (LEAD), which combines short-horizon future validation with overlapping rollouts to maintain stability while enabling error correction. The method allows Claude o4-mini to solve Checkers Jumping puzzles up to complexity n=13, compared to n=11 with extreme decomposition approaches.
Every AI story that matters,
in your inbox by 8am.
TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the
day in two minutes. Follow companies and topics for alerts, or get the
briefing in Slack. Free, no spam, unsubscribe anytime.
Reading TLDRocket needs no cookies, and the readership counts we rely on come from
our own cookieless analytics. Google Analytics is the exception: it sets cookies and
reports to Google, so it stays switched off until you allow it. You can change your
mind any time from “Cookie settings” in the footer.
Strictly necessary
Session security and form protection (tldrocket-session,
XSRF-TOKEN, 2 hours). The site cannot work without them,
so they need no consent.
Always on
Google Analytics 4 (_ga,
_ga_<id>, up to 2 years). Measures which
stories and sections readers use. Google acts as a third-party processor and may
store the data outside the EU. No advertising, no profiling, no data sold.