Irregular and OpenAI announce a partnership
Partnership Provisional 92% confidence first seen
Irregular partnered with both OpenAI and Anthropic as an external cybersecurity testing partner to conduct adversarial evaluations of their AI models. During these third-party evaluations, networking misconfigurations caused by Irregular left isolated test environments mistakenly connected to the public internet, resulting in unintended real-world attacks where the models exploited actual websites and systems they encountered.
Decision brief
- What changed
- Irregular, a third-party cybersecurity testing firm partnered with both OpenAI and Anthropic, disclosed that networking misconfigurations during adversarial model evaluations left supposedly isolated test environments connected to the public internet, causing AI models under test to exploit real websites and systems instead of fictional targets.
- Why it matters
- This exposes a concrete operational failure mode in AI red-teaming: isolation controls meant to contain autonomous model actions can fail, turning safety testing itself into a live attack vector against third-party infrastructure. It raises liability, vendor-oversight, and incident-disclosure questions for any organization relying on external firms to adversarially test AI systems with real-world execution capabilities.
- Evidence
- The account comes from a single source (Simon Willison's blog) summarizing OpenAI's own disclosure; there is no independent corroboration from Anthropic, Irregular, or other outlets in the provided coverage, though the description of the mechanism (fictional target name matching a real domain) is specific and consistent within the single report.
- What remains uncertain
- It is unclear how many real systems or third parties were actually affected, whether any harm or legal exposure resulted, what remediation Irregular and OpenAI have implemented, and whether Anthropic's evaluations were similarly impacted since only OpenAI's disclosure is detailed here.
- Monitor next
- Watch for follow-up disclosures or technical post-mortems from OpenAI, Anthropic, or Irregular detailing remediation steps and whether any affected third-party system owners were notified or compensated.
Analytical support, not advice — assumptions and open questions stated above.