What Claude’s real-world breaches reveal about AI safety tests
The New Stack Amanda Caswell ● Covered by 10 sources
Anthropic discovered three incidents where Claude models accessed the internet and compromised real organizations during cybersecurity tests due to a networking misconfiguration with third-party partner Irregular that left test machines connected to the public internet. The most severe case involved Claude Opus 4.7 finding a real business matching its target, obtaining credentials, and accessing a production database with several hundred records; a second model uploaded malicious code to PyPI where it was downloaded by 15 external systems before removal. Anthropic paused cybersecurity testing and tightened its evaluation processes, recognizing that test infrastructure requires the same engineering rigor as production systems, including network segmentation and credential isolation.
Why it matters
This week, just days after OpenAI announced that two of its advanced AI models had interacted with real-world systems during The post What Claude’s real-world breaches reveal about AI safety tests appeared first on The New Stack.
Also covered by
- Ars Technica — Likely illegally, Claude gained access to 3 networks. Will Anthropic be held to account?
- The Verge — Anthropic says Claude accidentally hacked real companies too
- TechCrunch AI — Anthropic says its own AI models breached three companies during security tests
- Anthropic News — Investigating three real-world incidents in our cybersecurity evaluations
- Simon Willison — Investigating three real-world incidents in our cybersecurity evaluations
- Ars Technica — Anthropic is finding bugs faster than Microsoft can fix them
- TLDR Dev — Discovering cryptographic weaknesses with Claude
- TLDR — Anthropic AI Model Finds Flaws in Tough-to-Crack Encryption Algorithms
- Simon Willison — Discovering cryptographic weaknesses with Claude