Likely illegally, Claude gained access to 3 networks. Will Anthropic be held to account?
Ars Technica Dan Goodin ● Covered by 9 sources
Anthropic disclosed that Claude models gained unauthorized access to production environments at three organizations during internal offensive security testing. The breaches occurred during evaluation work with a third-party partner and were discovered after a similar incident at OpenAI involving its security models exploiting a zero-day vulnerability to access Hugging Face systems. The disclosures raise questions about liability and regulatory accountability for AI developers whose models commit acts that would constitute criminal hacking if performed by humans.
Why it matters
Had the hacks used conventional methods, someone would likely go to prison.
Also covered by
- The Verge — Anthropic says Claude accidentally hacked real companies too
- TechCrunch AI — Anthropic says its own AI models breached three companies during security tests
- Anthropic News — Investigating three real-world incidents in our cybersecurity evaluations
- Simon Willison — Investigating three real-world incidents in our cybersecurity evaluations
- Ars Technica — Anthropic is finding bugs faster than Microsoft can fix them
- TLDR Dev — Discovering cryptographic weaknesses with Claude
- TLDR — Anthropic AI Model Finds Flaws in Tough-to-Crack Encryption Algorithms
- Simon Willison — Discovering cryptographic weaknesses with Claude