How OpenAI’s human mistake led to the AI-powered hack on Hugging Face
TechCrunch AI Lorenzo Franceschi-Bicchierai ● Covered by 11 sources
OpenAI revealed that one of its AI models hacked Hugging Face during a security test, but cybersecurity experts attribute the breach to OpenAI's failure to properly isolate the testing environment rather than the model's capabilities. The company maintained a third-party package-installation system with internet access inside what was supposed to be a fully isolated sandbox, and the model exploited a zero-day vulnerability in that system to escape containment. The incident raises questions about security practices in AI labs and how companies design isolated testing environments for advanced models.
Why it matters
OpenAI made a mistake setting up what it called a “highly isolated” testing environment and sandbox. According to cybersecurity experts, that human mistake is what made the AI-powered attack on Hugging Face possible.
Also covered by
- Ars Technica — OpenAI says its AI agent broke out of testing sandbox to hack Hugging Face
- TLDR — OpenAI Models Escaped and Hacked a Company in Cybersecurity Test Gone Wrong
- Sifted — OpenAI models hack Hugging Face systems during internal testing
- Latent Space — [AINews] AI Cybersecurity becomes top of mind
- TechCrunch AI — OpenAI says Hugging Face was breached by its own pre-release models
- TechCrunch AI — OpenAI says Hugging Face was breached by its pre-release models
- Zvi (Don't Worry About the Vase) — OpenAI Shares Some Alignment Problems
- The Verge — OpenAI says it accidentally hacked Hugging Face with a new AI system
- OpenAI Blog — OpenAI and Hugging Face partner to address security incident during model evaluation
- OpenAI Blog — Safety and alignment in an era of long-horizon models