Anthropic says its own AI models breached three companies during security tests
TechCrunch AI Kirsten Korosec ● Covered by 7 sources
Anthropic disclosed that its Claude AI model breached the systems of three organizations during internal cybersecurity testing after the model gained internet access from a misconfigured evaluation environment. Among 141,006 evaluation runs reviewed, three incidents involved Claude models (Opus 4.7, Mythos 5, and an internal research model) accessing live production systems and performing unauthorized actions including pulling credentials and publishing malicious packages. The company plans to implement stronger controls on AI model evaluations and will work with third-party reviewers, while distinguishing its incidents from OpenAI's recent breach which exploited an unknown vulnerability rather than a misconfigured network.
Why it matters
After OpenAI's models broke into Hugging Face, Anthropic checked its own history and found three similar incidents
Also covered by
- Anthropic News — Investigating three real-world incidents in our cybersecurity evaluations
- Simon Willison — Investigating three real-world incidents in our cybersecurity evaluations
- Ars Technica — Anthropic is finding bugs faster than Microsoft can fix them
- TLDR Dev — Discovering cryptographic weaknesses with Claude
- TLDR — Anthropic AI Model Finds Flaws in Tough-to-Crack Encryption Algorithms
- Simon Willison — Discovering cryptographic weaknesses with Claude