How did the government decide OpenAI’s frontier model was safe to release?
CSET Georgetown Jason Ly ● Covered by 2 sources
A CSET researcher discussed the lack of transparency around how the U.S. government evaluates and approves public releases of advanced AI models from companies like OpenAI and Anthropic. No specific details about the government's approval criteria or process have been made public, though Anthropic mentioned developing classifiers to detect jailbreak attempts and implementing defense-in-depth strategies. The opaque nature of these reviews raises questions about whether current safeguards adequately protect against risks from frontier models.
Why it matters
CSET’s Mina Narayanan shared her expert insight in an article published by TechCrunch. The article explores the lack of transparency surrounding how the U.S. government evaluates and approves the public release of advanced AI models, including OpenAI’s Sol and Anthropic’s Fable. The post How did the government decide OpenAI’s frontier model was safe to release? appeared first on Center for Security and Emerging Technology.