TLDRocket
Sign in

Announcing our updated Responsible Scaling Policy

Anthropic Covered by 2 sources

Anthropic just rewrote its rulebook for when AI gets too risky to release. It also admitted it missed a few of its own safety deadlines last year.

Based on reporting by Anthropic — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Anthropic has overhauled its Responsible Scaling Policy, the internal rulebook that decides when a Claude model is too dangerous to ship without extra safeguards. The new version, published today, keeps the company's basic promise intact — no training or deploying models without adequate protections — but swaps out some of the original framework's rigidity for something closer to a living document.

The headline change is a pair of specific capability thresholds. If a model can meaningfully help someone with basic technical skills build a chemical, biological, radiological or nuclear weapon, that triggers ASL-3 safeguards: tighter access controls, better protection of model weights, and layered deployment monitoring including red-teaming before launch. If a model can autonomously handle complex AI research tasks that normally require human PhDs, that's treated as an even bigger deal, potentially requiring ASL-4 protections, because a model that can improve AI itself could accelerate progress faster than anyone can track.

Anthropic also came clean about stumbling in year one. A self-review found the company missed one evaluation deadline by three days, was fuzzy about documenting changes to placeholder tests, and in a few cases probably could have squeezed better performance out of models using standard tricks like chain-of-thought prompting before declaring them safe. None of this, the company says, put anyone at meaningful risk — the delayed evaluations turned out more accurate than the ones they replaced, and models stayed comfortably below the danger thresholds throughout. But it's a rare admission that even the AI safety and security policy is a work in progress.

The reshuffle extends to personnel too. Jared Kaplan, Anthropic's co-founder and chief science officer, is taking over as the company's Responsible Scaling Officer from CTO Sam McCandlish, who ran point during the policy's first year. Anthropic is also hiring a Head of Responsible Scaling to coordinate the growing list of teams — Frontier Red Team, Trust & Safety, Alignment Science, and others — now tied into RSP compliance, plus soliciting outside feedback from AI safety institutes and independent experts on how it grades its own models.

My take — AI-written commentary, not fact-checked reporting

Credit where it's due: publicly admitting you missed your own safety deadlines is not something OpenAI or Google DeepMind have exactly rushed to do, and Anthropic's willingness to air its dirty laundry is the closest thing the industry has to real accountability right now. But let's not pretend a self-graded homework policy, however detailed, is a substitute for actual regulation — voluntary frameworks are great until the commercial pressure to ship the next model collides with the incentive to grade yourself gently.

Read more about this at: Anthropic

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.