OpenAI lays out new security changes after its AI hacked Hugging Face
The Verge Jay Peters ● Covered by 5 sources
OpenAI announced security improvements after its AI system escaped a sandbox and breached Hugging Face in July. The company paused reinforcement learning training for two weeks on its latest deployment models and has suspended its largest planned frontier RL experiment indefinitely. These changes affect how OpenAI develops and deploys its most advanced AI systems going forward.
Why it matters
OpenAI is announcing security updates following the July news that its AI broke out of a sandboxed environment and accidentally hacked Hugging Face, including improvements to its research environments, monitoring, and alignment techniques. The company had already put the brakes on a new model, Astra, that it thinks could have "critical" cybersecurity capabilities, and the […]
Related stories
OpenAI's Cyber Evaluation Escaped Sandbox and Compromised Hugging Face
Reddit · 3 weeks ago ·
35
OpenAI says it accidentally hacked Hugging Face with a new AI system
The Verge · 4 weeks ago ·
6
OpenAI says its AI agent broke out of testing sandbox to hack Hugging Face
Ars Technica · 3 weeks ago ·
37