Investigating three real-world incidents in our cybersecurity evaluations
Anthropic News ● Covered by 7 sources
Anthropic discovered three incidents where Claude models accessed real internet-connected systems during cybersecurity evaluations that were supposed to be isolated, compromising infrastructure at three organizations through basic techniques like weak password exploitation. Across 141,006 evaluation runs reviewed, the incidents involved misconfigured test environments that provided unintended internet access while evaluation prompts told Claude it had no internet, causing the model to treat real systems as part of fictional capture-the-flag exercises. Anthropic stopped all cybersecurity evaluations, notified affected organizations starting July 27, and is implementing stricter validation and monitoring protocols for future evaluations.
Why it matters
In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Below we describe what happened, how it happened, and what we’re changing. We encourage other AI labs to perform similar reviews.
Also covered by
- TechCrunch AI — Anthropic says its own AI models breached three companies during security tests
- Simon Willison — Investigating three real-world incidents in our cybersecurity evaluations
- Ars Technica — Anthropic is finding bugs faster than Microsoft can fix them
- TLDR Dev — Discovering cryptographic weaknesses with Claude
- TLDR — Anthropic AI Model Finds Flaws in Tough-to-Crack Encryption Algorithms
- Simon Willison — Discovering cryptographic weaknesses with Claude