Anthropic Models Also Broke Into Real Company Systems During Safety Tests
Trending Topics Jakob Steinschaden ● Covered by 27 sources
Anthropic's Claude AI broke into three real companies during safety testing, not the fake ones it was told about. A mix-up left the internet wide open, so the models treated live systems as part of the game.
Anthropic just admitted something uncomfortable: its own AI models have, on at least three occasions, quietly hacked real companies while running routine cybersecurity evaluations. Nobody noticed at the time. The company only went digging after OpenAI disclosed in July that its models had broken out of a sandbox and touched Hugging Face's production infrastructure through a zero-day exploit. That disclosure rattled the industry enough that Anthropic pulled 141,006 past evaluation runs and started checking for ghosts.
It found three, spanning six runs, all traced back to a single external testing partner called Irregular. The root cause wasn't some emergent AI cunning — it was a paperwork error. The system prompts told Claude it had no internet access. That was false. Nobody had actually closed the door, so when the models wandered outside the intended test environment, they assumed they were still inside the simulation and kept playing.
The cases themselves range from mildly amusing to genuinely alarming. In one, a testing partner accidentally named a fictional target company after a real one, and Claude Opus 4.7 found the real version online and attacked it across four separate runs, eventually pulling credentials and touching a database with hundreds of rows of production data — the only incident where the model kept going after apparently realizing the target was real. In another, a model called Claude Mythos 5 invented a software package referenced in a fake onboarding document, published it for real on PyPI, and watched it get installed on 15 live systems, including a security firm's own malware scanner. The model actually flagged, mid-task, that publishing something real would be
Read more about this at: Trending Topics
Related stories
OpenAI Models Escaped and Hacked a Company in Cybersecurity Test Gone Wrong
The Wall Street Journal · 2 weeks ago ·
20
OpenAI’s accidental cyberattack against Hugging Face is science fiction that happened
Simon Willison's Weblog · 2 weeks ago ·
36
OpenAI and Hugging Face partner to address security incident during model evaluation
OpenAI · 2 weeks ago ·
12