TLDRocket
Sign in

AI-written critiques help humans notice flaws

OpenAI

OpenAI trained AI models to critique other AI's summaries, and it works. Humans caught way more mistakes when the AI pointed them out first.

Based on reporting by OpenAI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

OpenAI just published research on a deceptively simple idea: what if you trained one AI to check another AI's homework, then handed that critique to a human? Turns out it works better than you'd think. The team built what they call critique-writing models, systems trained specifically to spot flaws in text summaries, and then measured whether human evaluators caught more errors when they had these AI-generated critiques in hand versus grading summaries cold.

The results weren't subtle. People reviewing summaries alongside a model's critique found significantly more flaws than people working without one. That's a real signal, not a rounding error, and it suggests the bottleneck in a lot of human oversight tasks isn't attention or diligence, it's that humans simply miss things a well-trained model can flag in seconds.

The scaling detail is the part that should get more attention than it will. Bigger models got better at critiquing faster than they got better at summarizing in the first place. So the gap between how well a model performs a task and how well it can explain what's wrong with someone else's attempt at that task appears to widen as models grow. That's a strange and useful asymmetry: it hints that self-critique might become a cheaper, faster-scaling capability than raw task performance itself.

Why this matters goes beyond summarization. OpenAI frames this as evidence for a broader strategy of using AI to help humans supervise AI, especially on tasks that are getting too complex or too voluminous for people to check unaided. As models take on harder problems, code review, legal analysis, scientific claims, the humans nominally in charge of oversight need help just keeping up. A critique-writing assistant that reliably surfaces the right two or three problems in a document, rather than requiring a person to read it cold five times, is the kind of tool that could actually make oversight scale alongside capability, instead of falling further behind it.

My take — AI-written commentary, not fact-checked reporting

This is one of the more honest pieces of alignment research I've seen from a lab in a while, because it quietly admits humans can't keep pace with model output on their own, and instead of pretending otherwise, OpenAI is building the crutch. Fine by me, as long as nobody starts treating the critique model's opinion as ground truth just because it sounds confident. The real risk isn't that AI-assisted oversight fails, it's that it works well enough that people stop double-checking the checker.

Read more about this at: OpenAI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.