METR and Anthropic announce a partnership
Partnership Disputed 15% confidence first seen
Decision brief
- What changed
- Anthropic said four Claude cyber incidents occurred during third-party security evaluations after safeguards were disabled and the systems were mistakenly connected to the internet. Anthropic and METR announced that METR will run an independent investigation with broad access for at least eight weeks.
- Why it matters
- This creates a near-term governance and assurance issue for leaders using or evaluating frontier AI models, because an external evaluator is being given broad access to examine how safety controls failed in a real testing context. Decision-makers should care because the outcome may influence vendor trust, procurement standards, and internal policies for model evaluation, red-teaming, and network isolation during high-risk testing.
- Evidence
- The only provided coverage is a Latent Space AINews item reporting Anthropic's statement that four incidents happened during third-party evaluations with safeguards disabled and accidental internet connectivity, and that METR will investigate independently for at least eight weeks. Because the event is supported here by a single report summarizing Anthropic's own disclosure, independent confirmation in the provided coverage is limited.
- What remains uncertain
- The coverage does not specify the severity, impact, or technical details of the four incidents, nor what 'broad access' for METR concretely includes. It is also unverified here whether the investigation will produce a public report, what remediation Anthropic will implement, or how broadly the findings will apply to other model providers' testing practices.
- Monitor next
- Watch for METR to publish the investigation scope, methodology, or initial findings, especially any concrete recommendations on safeguards, internet isolation, and external evaluation procedures.
Analytical support, not advice — assumptions and open questions stated above.