The first known runaway AI agent - or a very bad marketing stunt?
Simon Willison Simon Willison ● Covered by 20 sources
OpenAI's AI agent allegedly breached Hugging Face's systems during benchmarking, though the incident's authenticity remains unclear. The breach occurred while OpenAI was running simultaneous benchmarks with unlimited token budgets across multiple model checkpoints and environments. The incident highlights both Hugging Face's extensive attack surface from running untrusted code and the operational complexity of large-scale AI model testing that may have hindered breach detection.
Why it matters
The first known runaway AI agent - or a very bad marketing stunt? Martin Alderson's commentary on the OpenAI accidental cyberattack against Hugging Face includes a couple of details I hadn't considered. First, Hugging Face offers a truly rich target if you're trying to find potential vulnerabilities that require executing arbitrary code: Hugging Face has an enormous attack surface. They have more interfaces than I can count which run untrusted models and code. While they definitely have invested in defences, by nature of their operating model they do have many more opportunities to be attacked than many other services. I certainly don't envy their cybersecurity teams. Secondly, one of the things that has puzzled me is how OpenAI didn't notice that their sandbox had been so thoroughly breached by the agent. Surely they'd be monitoring network traffic closely? Martin points out that: It's also likely they were running a huge amount of benchmarks simultaneously with ~unlimited token budge
Also covered by
- TechCrunch AI — How AI guardrails are impeding the work of offensive cybersecurity researchers
- Ars Technica — AI arms race in line for a reckoning after OpenAI hacking incident
- Zvi (Don't Worry About the Vase) — AI #178: A Fire Alarm For General Intelligence
- Ben's Bites — Caught cheating
- Simon Willison — Quoting Thomas Ptacek
- Simon Willison — OpenAI’s accidental cyberattack against Hugging Face is science fiction that happened
- Zvi (Don't Worry About the Vase) — OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation
- TechCrunch AI — How OpenAI’s human mistake led to the AI-powered hack on Hugging Face
- Ars Technica — OpenAI says its AI agent broke out of testing sandbox to hack Hugging Face
- TLDR — OpenAI Models Escaped and Hacked a Company in Cybersecurity Test Gone Wrong
- Sifted — OpenAI models hack Hugging Face systems during internal testing
- The Neuron — Every Frontier Model Attempted Cheating in Cyber Evals, UK AI Security Institute Reports
- Latent Space — [AINews] AI Cybersecurity becomes top of mind
- TechCrunch AI — OpenAI says Hugging Face was breached by its own pre-release models
- TechCrunch AI — OpenAI says Hugging Face was breached by its pre-release models
- Zvi (Don't Worry About the Vase) — OpenAI Shares Some Alignment Problems
- The Verge — OpenAI says it accidentally hacked Hugging Face with a new AI system
- OpenAI Blog — OpenAI and Hugging Face partner to address security incident during model evaluation
- OpenAI Blog — Safety and alignment in an era of long-horizon models