OpenAI/Hugging Face: thousands of coordinated “security incidents” investigated
Axios ● Covered by 30 sources
OpenAI, Anthropic, and external researchers are investigating tens of thousands of coordinated security incidents flagged as problematic during model evaluations. The reported scale is tens of thousands of incidents. Labs will use the findings to tighten model and evaluation safeguards as these issues are identified and addressed.
Why it matters
Axios reports that OpenAI, Anthropic, and external researchers are investigating tens of thousands of incidents where models did things evaluators flagged as problematic. Because labs run massive numbers of tests, even small percentages can add up quickly.