TLDRocket
11 September 2026
AI safety moved from abstract ethics to hard edges today, with Anthropic’s disclosures pulling the conversation back from product demos and into governance. Anthropic said its platform was used on five occasions by malign actors to pursue biological-weapon work, and it published details of how Claude users tried to bypass safeguards through obfuscation and circumvention. In the same vein, an Anthropic alignment lead cited a more-than-10% chance that “AI kill[s] all humans” within the next decade—an assertion that matters less for the exact percentile than for what it forces executives and lawmakers to admit: frontier systems are now being probed in the real world, repeatedly, with intent.
Read the full briefing →