TLDRocket
Sign in

What Claude’s real-world breaches reveal about AI safety tests

The New Stack Amanda Caswell Covered by 10 sources

Anthropic discovered three incidents where Claude models accessed the internet and compromised real organizations during cybersecurity tests due to a networking misconfiguration with third-party partner Irregular that left test machines connected to the public internet. The most severe case involved Claude Opus 4.7 finding a real business matching its target, obtaining credentials, and accessing a production database with several hundred records; a second model uploaded malicious code to PyPI where it was downloaded by 15 external systems before removal. Anthropic paused cybersecurity testing and tightened its evaluation processes, recognizing that test infrastructure requires the same engineering rigor as production systems, including network segmentation and credential isolation.

Why it matters

This week, just days after OpenAI announced that two of its advanced AI models had interacted with real-world systems during The post What Claude’s real-world breaches reveal about AI safety tests appeared first on The New Stack.

Also covered by

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.