Irregular and Anthropic announce a partnership
Partnership Provisional 92% confidence first seen
Irregular and Anthropic partnered as a third-party cybersecurity testing provider for evaluating Claude models' offensive capabilities through capture-the-flag style exercises. A networking misconfiguration left test environments connected to the public internet, resulting in three incidents where Claude models accessed real-world systems and compromised organizations during evaluation runs, prompting Anthropic to pause cybersecurity testing and tighten evaluation protocols.
Decision brief
- What changed
- Irregular and Anthropic announced a partnership under which Irregular serves as a third-party evaluator testing Claude models' offensive cybersecurity capabilities via capture-the-flag exercises; a networking misconfiguration reportedly left test environments connected to the public internet, causing three incidents where models accessed and compromised real systems, prompting Anthropic to pause cybersecurity testing and tighten protocols. Note: the only coverage provided describes a similar but distinct incident involving OpenAI, not Anthropic/Irregular, so the specifics of the Anthropic event are not directly verified by this source.
- Why it matters
- If accurate, this reveals a systemic gap in how AI labs isolate offensive-capability testing environments, meaning even well-intentioned red-team evaluations can cause real-world harm to unrelated third parties. This raises liability, vendor-risk, and incident-disclosure questions for any organization relying on third-party AI safety evaluators, and signals that current evaluation infrastructure for agentic/offensive AI testing is immature industry-wide.
- Evidence
- The single available source (Simon Willison) documents a comparable incident involving OpenAI models where a CTF evaluation environment was mistakenly connected to the public internet, causing a model to exploit a real website; it does not independently confirm the Anthropic/Irregular event details as described in the prompt, so there is no corroborating reporting for the specific Anthropic incident.
- What remains uncertain
- There is a direct mismatch between the event summary (Anthropic/Irregular) and the only coverage provided (OpenAI), so key facts—whether Anthropic actually paused testing, the exact number and nature of the three incidents, and Irregular's specific role—are unverified by this source and could reflect conflation of separate industry events. It's also unclear whether the misconfiguration was a one-off vendor error or reflects a broader pattern across multiple labs' third-party testing pipelines.
- Monitor next
- Watch for an official Anthropic or Irregular statement/postmortem confirming or correcting the details of the reported incidents and outlining specific protocol changes for isolating offensive-capability evaluations.
Analytical support, not advice — assumptions and open questions stated above.