TLDRocket
Sign in

Finding GPT-4’s mistakes with GPT-4

OpenAI

OpenAI built CriticGPT, an AI that reads ChatGPT's answers and points out what's wrong with them. It's basically hiring a robot editor because human reviewers keep missing bugs in AI-written code.

Based on reporting by OpenAI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

OpenAI has a problem that gets harder as its models get smarter: the humans grading ChatGPT's answers during training increasingly can't tell when the AI is wrong. Code is the clearest example. A response can look clean, run without errors, and still contain a subtle logic flaw that only someone who wrote similar code for years would catch. So OpenAI built CriticGPT, a model derived from GPT-4, whose entire job is to write critiques of ChatGPT's output and flag the mistakes a busy human trainer might scroll right past.

The setup is straightforward in concept, messier in practice. OpenAI trained CriticGPT on a pile of ChatGPT responses that contained deliberately inserted bugs, then had it learn to describe what was broken and why. In head-to-head tests, human reviewers armed with CriticGPT's critiques caught more real problems in code than reviewers working alone, and they caught more than reviewers paired with other ChatGPT models playing critic. That's the whole pitch: not a replacement for human judgment, but a second pair of eyes that doesn't get tired or distracted.

OpenAI is candid about where this breaks down. CriticGPT was trained on relatively short, self-contained answers, so its usefulness on sprawling, multi-step tasks is unproven. It can also hallucinate its own criticisms, inventing problems that aren't there, which forces trainers to double-check the critic the same way they'd double-check the model being critiqued. And when a task involves errors spread across many parts of a long answer, CriticGPT tends to miss the cumulative picture, since it was tuned to flag concentrated, well-defined bugs rather than diffuse ones.

What makes this notable isn't the tool itself so much as the timing. OpenAI is essentially admitting that RLHF, the human-feedback method that shaped ChatGPT's behavior, is running into a ceiling: models are outpacing the reviewers meant to police them. CriticGPT is a stopgap for that mismatch, a way to keep human oversight relevant a little longer by giving humans AI-generated leverage. It's a tacit acknowledgment that scaling these systems responsibly requires scaling the tools for checking them just as fast, if not faster.

My take — AI-written commentary, not fact-checked reporting

This is OpenAI quietly conceding that human oversight of frontier models is already strained, and that's the real story buried under the CriticGPT branding. Using an AI to catch an AI's mistakes is a reasonable stopgap, but it's also a recursive patch job that only holds until CriticGPT itself needs a critic. I'd rather see labs invest as much in transparency about failure rates as they do in the tooling that papers over them.

Read more about this at: OpenAI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.