What Actually Keeps an AI Benchmark Useful? Scale
StackSweep
A new study says nearly half of the AI benchmarks everyone cites are basically maxed out. Top models score so close together the rankings are just noise now, and secrecy or format tricks don't fix it.
There's a quiet crisis brewing in how the AI industry measures itself, and a new paper from the EvalEval Coalition just put a number on it. Out of 60 widely used text benchmarks, 29 are so saturated that top models can no longer be meaningfully separated on them. Fourteen have crossed a threshold where the leaderboard is, statistically speaking, mostly noise. That means the "state of the art" headline you read last week, built on a 0.3-point gap between two models, might not reflect any real difference in capability at all.
The study's real contribution is defining saturation rigorously instead of just gesturing at it. The researchers built a Saturation Index comparing the score spread among the top five models against the expected statistical noise for that benchmark's test set size. Score zero means wide open, score one means the leaderboard tells you nothing. They ran this across 60 benchmarks aged between 1 and 114 months, with 23 researchers hand-annotating 14 properties each, and the pattern is stark: mean saturation climbs from 0.51 in benchmarks under two years old to 0.60 in ones over five years old, with the share of highly saturated benchmarks nearly doubling.
What's more interesting is what turned out not to matter. The field has long assumed private test sets protect against gaming and staleness, but the paper found no significant gap between public and private benchmarks. Multiple-choice formats, supposedly harder to game than open-ended ones, showed no advantage either. Multilingual benchmarks looked more durable at first glance, until the researchers noticed they're simply younger on average, by 16 months, and that age gap explains the effect entirely. Even citation count, often used as a proxy for a benchmark's continued relevance, loses its statistical significance once age gets controlled for.
So what actually predicts durability? Test set size, mostly, along with expert curation, though the latter is tangled up with its own age confound since crowdsourced datasets in the sample tend to be older. A combined Bayesian model using all the tested factors together explained 88.4% of the variance in saturation, and age plus test set scale did nearly all the work. That's a fairly blunt finding for a field that spends a lot of energy debating benchmark design choices that, per this analysis, barely move the needle compared to simply making the test set bigger.
The practical fallout is that leaderboards need uncertainty bars, not just point scores, and benchmarks need retirement plans instead of infinite shelf lives built on accumulated citations. None of that requires new evaluation science. It requires someone to actually report the confidence intervals and admit when a benchmark has run out of runway, which is a cultural problem more than a technical one.
My take
None of this should surprise anyone who's watched a benchmark get quietly retired the moment it stops flattering the incumbent labs' models. The real scandal here isn't that benchmarks saturate, it's that leaderboards keep reporting them as gospel without error bars, letting marketing teams cite noise as a win. Bigger test sets and expert curation aren't glamorous fixes, but they're cheap compared to the credibility cost of another "new SOTA" claim that's actually a coin flip.
Read more about this at: StackSweep