TLDRocket
Sign in

Introducing SimpleQA

OpenAI

OpenAI just dropped SimpleQA, a new test for catching how often chatbots confidently make stuff up. It's short factual questions with one right answer, so there's nowhere to hide.

Based on reporting by OpenAI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

OpenAI has a new yardstick for one of AI's most stubborn problems: models that state wrong facts with total confidence. SimpleQA is a benchmark built entirely around short, fact-seeking questions that each have a single, verifiable answer, the kind of thing you'd expect a competent research assistant to just know or admit not knowing.

What makes this different from the usual trivia-style benchmarks is the design intent. The questions were specifically chosen to be hard for frontier models like GPT-4, not easy softballs that every system aces immediately, which would make the benchmark useless within a year. Each question also went through human verification, so there's a clean, unambiguous ground truth to grade against instead of fuzzy partial-credit scoring.

Grading itself is automated, sorting responses into three buckets: correct, incorrect, or not attempted. That last category matters more than it might seem. A model that says 'I don't know' when it genuinely doesn't know is behaving better than one that guesses wrong and states it as fact. OpenAI is explicitly using SimpleQA to look at calibration, whether a model's confidence actually lines up with how often it's right, which is really a proxy for measuring hallucination rates in a controlled way.

This lands at a moment when factual reliability is arguably the biggest gap between AI marketing and AI reality. Bigger models keep getting better at reasoning and coding benchmarks, but a system that invents a plausible-sounding but false fact still causes real damage, especially in things like search, research, or customer support where users can't easily fact-check every claim. A benchmark that isolates just this failure mode, stripped of complex reasoning or multi-step tasks, gives researchers a cleaner signal to actually improve on.

My take — AI-written commentary, not fact-checked reporting

Finally, a benchmark that isn't just chasing bigger reasoning scores and instead asks the boring but crucial question: can this thing tell the truth about simple facts. I'd rather see labs compete on hallucination rates than on leaderboard flexing, because that's the metric that actually determines whether I'd trust a model's answer without double-checking it myself.

Read more about this at: OpenAI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.