AI-written critiques help humans notice flaws
OpenAI Blog
Researchers trained AI models to write critiques identifying flaws in AI-generated summaries. Human evaluators detected flaws in summaries 30-40% more often when shown the model critiques compared to evaluating summaries alone. The approach suggests AI systems could help humans supervise other AI systems on complex evaluation tasks.
Why it matters
We trained “critique-writing” models to describe flaws in summaries. Human evaluators find flaws in summaries much more often when shown our model’s critiques. Larger models are better at self-critiquing, with scale improving critique-writing more than summary-writing. This shows promise for using AI systems to assist human supervision of AI systems on difficult tasks.