OpenAI and Anthropic share findings from a joint safety evaluation
OpenAI Blog
OpenAI and Anthropic released results from a joint safety evaluation where they tested each other's AI models across multiple failure modes including misalignment, instruction-following errors, hallucinations, and jailbreak vulnerabilities. The evaluation assessed both companies' most capable models available at the time of testing, with results published to document specific strengths and weaknesses in each system. The findings demonstrate that direct collaboration between competing labs can identify safety issues more comprehensively than single-lab evaluations.
Why it matters
OpenAI and Anthropic share findings from a first-of-its-kind joint safety evaluation, testing each other’s models for misalignment, instruction following, hallucinations, jailbreaking, and more—highlighting progress, challenges, and the value of cross-lab collaboration.