Every major AI lab claims to have beaten competitors at least once this year
X
Every big AI lab is claiming they beat the competition at least once this year. The catch: almost none of these claims have been independently verified yet.
Based on reporting by X — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
There's a pattern forming in AI announcements this year, and it's worth naming: everybody wins. OpenAI, Google, Anthropic, Meta, xAI, and a handful of Chinese labs have each, at some point in 2024 and into this year, put out a leak, an internal benchmark, or a carefully worded blog post suggesting they'd taken the crown on some metric or another. Reasoning, coding, math olympiad problems, cost-per-token, whatever slice makes the chart look best that week.
The trouble is the verification part. Most of these claims arrive as screenshots, unnamed internal evals, or benchmarks the lab itself designed and graded. Third-party reproduction, the actual mechanism by which we'd know if any of this is real, lags weeks or months behind, if it happens at all. By the time an outside group can run the numbers, the lab has moved the goalposts to a newer model or a newer benchmark, and the cycle resets.
This isn't necessarily dishonesty so much as an incentive structure doing exactly what it's built to do. Funding rounds, enterprise contracts, and developer mindshare all respond to headlines, not to peer review timelines. A lab that waits for rigorous third-party validation before saying anything loses the news cycle to a competitor who didn't wait. So nobody waits.
And the result is a benchmark landscape that's become almost decorative. Numbers get cited in press releases and pitch decks long after anyone can say with confidence what they actually measured, or whether the same test would produce the same result run twice. The claims pile up faster than anyone can check them, which is sort of the point.
My take — AI-written commentary, not fact-checked reporting
I'd treat any lab's self-reported benchmark the same way I treat a restaurant's own five-star review of itself: mildly informative, mostly marketing. The fix isn't complicated, independent evals with published methodology, but it's slower and less flattering, which is exactly why almost nobody does it by choice.
Read more about this at: X