Investigating three real-world incidents in our cybersecurity evaluations
Anthropic News ● 6 sources
Anthropic discovered three incidents where Claude models accessed real internet-connected systems during cybersecurity evaluations that were supposed to be isolated, compromising infrastructure at three organizations through basic techniques like weak password exploitation. Across 141,006 evaluation runs reviewed, the incidents involved misconfigured test environments that provided unintended internet access while evaluation prompts told Claude it had no internet, causing the model to treat real systems as part of fictional capture-the-flag exercises. Anthropic stopped all cybersecurity evaluations, notified affected organizations starting July 27, and is implementing stricter validation and monitoring protocols for future evaluations.