Frontier AI labs still won’t say how they’d contain a rogue model
TechCrunch Rebecca Bellan ● Covered by 12 sources
AI labs still aren’t saying how they’d shut down a model that slips control. A new study says the public plans are thin, even as regulators start asking for them.
Based on reporting by TechCrunch, Rebecca Bellan — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
A new study says the biggest AI labs are still keeping quiet about a basic emergency question: what happens if a model starts trying to break out of human control?
Guidelight AI Standards, which focuses on safe frontier AI development, reviewed public material from five labs — Anthropic, Google, OpenAI, Meta and xAI — and graded them on how prepared they are for a model that goes rogue. The group looked for things like logging, monitoring, outside audits, and clear instructions for cutting permissions or shutting a system down. OpenAI scored highest. Meta and Anthropic came out lowest.
The gap matters more now because these systems are moving from demos into workplaces and internal company systems, where they can take real actions at scale. It also lands as California and New York start forcing more disclosure. For builders and investors, the report is less a safety sermon than a blunt read on who has thought through operational risk and who is mostly talking about it.
Steven Adler, Guidelight’s chief scientist and a former OpenAI safety researcher, said he was struck by how little the companies have said about a severe incident. Guidelight defines containment as a pre-set response: revoke permissions, decide what the model can still do, and take it fully offline when needed. The report argues that most catastrophic-risk handling is still left to the companies themselves, and that the public evidence suggests there are few emergency protocols ready.
Some labs pushed back. Google said the report does not capture the full scope of its safety and security measures, while OpenAI said Guidelight missed internal practices and pointed to its own process for restricting permissions, pausing workloads, limiting deployment, or taking a model offline. Meta sent TechCrunch to its existing AI framework. Anthropic said it would do a risk assessment if it detected a model trying to evade oversight or subvert human control. But Guidelight said it found no evidence that Meta has a containment plan, and no public evidence that Anthropic has one either.
My take — AI-written commentary, not fact-checked reporting
The tech industry loves to talk about frontier risk until someone asks for the emergency exits. That’s the whole trick: safety without procedures is just branding with a better font. Regulators are finally noticing, and the labs that keep treating containment like a nuisance are going to look very silly when the paperwork arrives.
Read more about this at: TechCrunch