What Actually Keeps an AI Benchmark Useful? Scale
StackSweep 3 weeks ago 39
Nearly half of 60 widely used language model benchmarks have become saturated, meaning top models score within statistical noise of each other and the benchmark can no longer rank them. Of the 60 benchmarks analyzed, 29 show high or very high saturation (Saturation Index ≥0.7), with saturation climbing from a mean of 0.51 for benchmarks under 24 months old to 0.60 for those over 60 months old. The study found that assumed safeguards like private test sets, harder output formats, and multilingual scope provide no meaningful protection once benchmark age is controlled for, leaving only test set scale and expert curation as effective strategies, and recommending explicit retirement criteria rather than letting benchmarks accumulate citations indefinitely.