Announcing our updated Responsible Scaling Policy
Anthropic News ● Covered by 2 sources
Anthropic updated its Responsible Scaling Policy, a risk governance framework for frontier AI systems, to introduce more flexible capability thresholds and refined safeguard assessment processes. The updated policy defines two key capability thresholds requiring upgraded safeguards: autonomous AI research and development capabilities, and meaningful assistance with creating chemical, biological, radiological, or nuclear weapons. Models reaching these thresholds will require enhanced security standards (ASL-3 or ASL-4) including internal access controls, deployment monitoring, and pre-deployment red teaming.
Why it matters
Today we are publishing a significant update to our Responsible Scaling Policy (RSP), the risk governance framework we use to mitigate potential catastrophic risks from frontier AI systems.