Finding GPT-4’s mistakes with GPT-4
OpenAI Blog
OpenAI developed CriticGPT, a GPT-4-based model, to identify errors in ChatGPT outputs and assist human trainers during reinforcement learning from human feedback. CriticGPT achieved 63% agreement with human trainers on whether responses contained mistakes, compared to 51% for humans evaluating other humans' critiques. The tool aims to reduce the labor required for human feedback collection in model training by automating initial error detection.
Why it matters
CriticGPT, a model based on GPT-4, writes critiques of ChatGPT responses to help human trainers spot mistakes during RLHF