Building better AI benchmarks: How many raters are enough?
Google Research
Google Research says most AI benchmarks use way too few human raters to be trustworthy. Turns out 3-5 people rating each example isn't nearly enough to capture real disagreement.
Based on reporting by Google Research — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Ask five people if a tweet is toxic and you might get five different answers. Ask five hundred, and patterns start to emerge — patterns that most AI benchmarks never bother to look for. That's the uncomfortable finding buried in new research out of Google, which spent months building a simulator to test one deceptively simple question: when you're paying humans to label data for AI evaluation, is it better to have more items or more raters per item?
The team ran what amounts to a massive stress test across four real-world datasets — a 107,620-comment toxicity corpus, the DICES chatbot safety set, the cross-cultural D3code dataset spanning 21 countries, and a job-tweets collection called Jobs. They dialed the number of rated items from 100 up to 50,000, and the number of raters per item from 1 all the way to 500, checking which combinations actually produced statistically reliable, reproducible results. The answer challenges a habit that's been baked into machine learning for years: using 1, 3, or 5 raters and calling it a day. According to this study, that's often not enough breadth or depth to mean anything. Ten raters per item, at minimum, is closer to the real threshold.
What's more interesting is that there's no single correct ratio — it depends entirely on what you're trying to measure. If you just want to know whether a model agrees with the majority opinion, spreading your budget across more items beats piling more raters onto each one. But if you care about capturing the full spread of human opinion — the gap between a firm "yes" and a hesitant "maybe" — then depth wins, and only adding more raters per item reveals that variation. Mix up the strategy and metric, and you can burn through a huge budget while still landing on unreliable conclusions.
There's a silver lining buried in the math, though. The researchers found that roughly 1,000 total annotations, allocated the right way, can produce solid, reproducible results. You don't need Google-scale budgets to do this well; you need to pick the right ratio for the right question. That's a genuinely useful, actionable takeaway for smaller labs and startups that can't afford to survey tens of thousands of people every time they want to check a model.
The bigger point here is about what AI evaluation has quietly assumed for a decade: that somewhere under all the noise, there's one true label waiting to be found. As AI gets pushed into judgment calls about tone, harm, and intent — areas where reasonable humans genuinely disagree — that assumption starts to crack. Google has open-sourced the simulator on GitHub, which means anyone can now test their own benchmarks against this breadth-versus-depth tradeoff instead of just guessing.
My take — AI-written commentary, not fact-checked reporting
I've long suspected that a lot of AI benchmark scores are dressed-up noise, and this confirms it — five raters per item was never a serious sample size, it was just what fit the budget spreadsheet. What bugs me is how long this went unexamined while entire leaderboards and papers built careers on numbers this shaky. Open-sourcing the simulator is the right move, and honestly every lab claiming a benchmark win should have to show their rater count before anyone takes the result seriously.
Read more about this at: Google Research