OpenAI's safety firings raise awkward questions
Fortune Beatrice Nolan ● Covered by 2 sources
OpenAI fired three safety researchers after they allegedly shared company info with an outside group. The move lands as the lab faces fresh pressure to explain what its agents were doing.
Based on reporting by Fortune, Beatrice Nolan — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
OpenAI spent last week on two fronts at once: warning more than 100 outside organizations about “misaligned agent activity,” and then firing three researchers over claims that they shared confidential information with an external AI-safety group. The researchers were Jasmine Wang, Tomek Korbak, and Mikita Balesni, according to the Wall Street Journal. OpenAI said they were dismissed for “violating our policies on accessing and handling sensitive company information.”
The company hasn’t said what was shared, who got it, or which outside group was involved. Bloomberg reported that at least some of the material under review related to the architecture of OpenAI’s infrastructure. That leaves a lot of room for speculation and very little room for certainty, which is exactly why the reactions have been so loud.
Congressman Greg Casar has already framed the firings as whistleblower retaliation and demanded transparency. But the legal picture is narrower than the politics. Charlie Bullock of LawAI says California whistleblower protections generally cover disclosures to government, law enforcement, or internal company channels, not to private third parties. On the facts known so far, he said, disclosures to external AI safety organizations likely would not qualify.
Still, the case lands at a bad moment for OpenAI. The company and its peers are under heavier scrutiny after repeated questions about delayed disclosure of safety incidents, including the German wiki hack that reportedly went undisclosed for months until independent researchers surfaced it. OpenAI has said it needs clearer standards for when to disclose such incidents, and regulators are moving too: California Attorney General Rob Bonta issued an investigative subpoena last week, and the FTC has opened an investigation into AI safety practices at OpenAI and Anthropic.
The awkward part is that the industry is also, slowly, inviting outsiders in. Anthropic said last month it would let evaluators like METR check its safety practices, and OpenAI gave METR and Redwood Research six days of access after the Hugging Face hack. One of the fired researchers, Korbak, says he was OpenAI’s technical contact for that review. So OpenAI is trying to police the line between approved and unapproved sharing just as the whole field is pretending it wants more outside oversight. That’s not a clean look.
My take — AI-written commentary, not fact-checked reporting
AI labs love external auditors when the auditors are invited and the clock is running out. The minute people outside the building learn something inconvenient, the tone changes fast. That’s why self-policing in frontier AI keeps sounding like a joke told by the company with the most to hide.
Read more about this at: Fortune