TLDRocket
Sign in

GPT-6 Still Behind Fable 5.1 As Artificial Analysis Overhauls Intelligence Index

Trending Topics Jakob Steinschaden Covered by 9 sources

Artificial Analysis rewrote its AI ranking, and Claude Fable 5.1 still sits on top. GPT-6 Astra improved, but OpenAI still trails Anthropic at the summit.

Based on reporting by Trending Topics, Jakob Steinschaden — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Artificial Analysis has overhauled its Intelligence Index, the score people use when they want to compare language models against the rest of the field. The new version, 4.2, is meant to be tougher and more realistic, with more hidden test data baked in. But the headline hasn’t changed: Anthropic’s Claude Fable 5.1 remains first, and OpenAI’s GPT-6 Astra stays in second.

The company is calling this an interim release. Eight months have passed since Index v4, and Artificial Analysis says it held back bigger changes to keep the benchmark steady through several major model launches. Now it’s pulling forward parts of the planned version 5 because the frontier has shifted so fast in recent weeks. More small updates are expected after this.

The benchmark itself has changed in a few important ways. AA-Briefcase, Artificial Analysis’ in-house test for realistic agentic knowledge work, joins the Index. So does GDP.pdf, a Surge AI benchmark that asks models to reason across 100 PDFs from ten domains, spanning 4,592 pages. At the same time, GPQA Diamond drops out because it has become too easy for frontier models to separate themselves on it. Private held-out test sets now make up 40 percent of the overall weighting, up from v4.1, and the company says that share will grow again in version 5.

The updated scoring also tightens the screws elsewhere. AA-LCR has moved to version 1.1 with corrected answer keys, GDPval-AA v2 and AA-Briefcase got better sampling and a re-anchored Elo scale, and SciCode now gives credit to code that is slow but correct. The point is simple enough: it should be harder for labs to tune to the benchmark instead of the work the benchmark is supposed to represent.

The rankings still tell a familiar story. Claude Fable 5.1 leads overall, with GPT-6 Astra four points ahead of GPT-5.6 Sol but still behind the Anthropic model. Meta sits third among labs, followed by SpaceXAI, Moonshot with Kimi, Z.AI and Google. On AA-Briefcase, Anthropic’s Fable 5.1 and Opus 5 are ahead of Astra. On GDP.pdf, though, OpenAI flips the script: Astra scores 33.2 percent, above GPT-5.6 Sol at 28.2 percent and Claude Fable 5.1 at 26.2 percent.

My take — AI-written commentary, not fact-checked reporting

This is what a serious benchmark looks like when it stops being a leaderboard toy and starts trying to resist gaming. The bad news for the industry is that the score got harder and the top spot still didn’t move. That’s the problem with AI hype: every model launch sounds like arrival, and then the benchmark gets adjusted and reality strolls back in through the side door.

Read more about this at: Trending Topics

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.