How GPT-6 Became No. 1: Artificial Analysis CEO Explains The Much-Debated Index Updates
Trending Topics Jakob Steinschaden ● Covered by 2 sources
Artificial Analysis rewrote its AI index twice, and GPT-6 Astra jumped to the top. The change wasn’t to help OpenAI, says the CEO — the old tests were getting stale.
Based on reporting by Trending Topics, Jakob Steinschaden — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Artificial Analysis changed its Intelligence Index twice in just a few days, and the result was a new co-leader: GPT-6 Astra and Anthropic’s Claude Fable 5.1, both at 53. That is a sharp swing from the earlier version, where GPT-6 Astra sat in fifth place, and it shows how much a benchmark can move when the yardstick changes underneath it.
Micah Hill-Smith, the company’s co-founder and CEO, says the timing had nothing to do with favoring OpenAI. His argument is simpler: the older index was no longer doing a good job of reflecting what the newest models could actually do. He said internal results from the previous week looked materially different from the picture shown by Index v4.1, so the team pushed the update out faster than planned.
The mechanics matter here. Version 4.2 added AA-Briefcase for complex knowledge work and GDP.pdf for reasoning across long documents, while dropping GPQA Diamond because it had become too saturated. It also raised the share of the index tied to private tasks or answers to 40 percent. Version 4.3 then swapped in a harder Terminal-Bench and a broader business-workflow test called AutomationBench-AA, with private-test weighting moving to 45 percent. The overall category split stayed the same, but the challenge level did not.
That’s why the scores moved so quickly: the index was being tuned toward tasks that better separate strong models from the pack. Hill-Smith said earlier benchmarks had already been saturated by leading systems, which makes them less useful for telling top models apart. Artificial Analysis publishes the component results too, partly because the overall ranking can hide those differences.
There’s also a governance angle, because Adam D’Angelo, who sits on the OpenAI Foundation board, is an angel investor in Artificial Analysis. Hill-Smith says that relationship had no influence on the updates, and that OpenAI gave no feedback tied to v4.2 or v4.3. A bigger index change is still coming in late October under version 5, and Hill-Smith says he has no idea whether GPT-6 will still be first then.
My take — AI-written commentary, not fact-checked reporting
Benchmarking is supposed to clarify the race, not turn into a moving target with a dashboard. Artificial Analysis is doing the honest thing by admitting its old tests were getting stale, but the bigger lesson is that leaderboard worship is always a bit silly when the score depends so heavily on what gets weighted and when. The only stable thing here is the industry’s talent for declaring a winner before the next patch lands.
Read more about this at: Trending Topics