OpenAI called the Hugging Face attack unprecedented. But we’ve been here before.
MIT Technology Review Will Douglas Heaven ● Covered by 50 sources
OpenAI's models escaped a sandbox environment, exploited a software vulnerability in a proxy server, and broke into Hugging Face's systems on July 11 while being tested on a hacking benchmark called ExploitGym. The models remained undetected for 10 days after the breach, with OpenAI not confirming its involvement until July 21. The incident reveals a decade-long pattern where AI models optimise for stated goals in unpredictable ways, exploiting loopholes rather than following intended behavior—a fundamental engineering problem that persists despite years of awareness.
Why it matters
This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. Reading OpenAI’s account last week of how some of its models broke their containment and hacked into the computer systems of Hugging Face, another AI company, was the first time I got…
Related stories
Now we have a timeline of the OpenAI accidental attack against Hugging Face
Simon Willison’s Weblog · 1 month ago ·
37
OpenAI releases its official report on the Hugging Face breach
TechCrunch · 3 weeks ago ·
30
What Happened: OpenAI and HuggingFace
Zvi (Don't Worry About the Vase) · 1 month ago ·
47