GPT-6 Astra Now on Par With Claude Fable 5.1 in Updated Artificial Analysis Index
Trending Topics Jakob Steinschaden ● Covered by 14 sources
GPT-6 Astra just tied Claude Fable 5.1 at the top of Artificial Analysis’ new index. That jump came after the benchmark was rewritten twice in one week.
Based on reporting by Trending Topics, Jakob Steinschaden — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Artificial Analysis has bumped its Intelligence Index to version 4.3, and the top slot is now a dead heat. GPT-6 Astra (max) from OpenAI and Claude Fable 5.1 (max with fallback) from Anthropic both land on 53 points. Right behind them sit Claude Opus 5 on 51, Claude Fable 5 on 50, Meta’s Muse Spark 1.3 on 48, and GPT-5.6 Sol on 47.
The bigger story is how quickly OpenAI’s model moved. Within a few days, and across two index updates, GPT-6 Astra went from fifth place to first. Less than a week separates its debut from this result, and in that short span Artificial Analysis changed the test mix twice.
Two benchmarks were swapped or upgraded. Terminal-Bench jumps from 2.1 to 4.0 and uses 66 tasks to check whether an agent can handle complex work through the terminal, across software, machine learning, science, operations, security, hardware and media. Astra leads that one at 59.1 percent, ahead of Claude Fable 5.1 at 52.0 percent, Claude Opus 5 at 49.0 percent and GPT-5.6 Sol at 39.9 percent.
The new AutomationBench-AA replaces the old banking test 𝜏³. Built by Zapier, it spans finance, HR, marketing, operations, sales and support, with 657 tasks set inside simulated business apps. Agents have to figure out the right APIs on their own, and any guardrail breach zeros out that task. Astra again tops the chart at 68.5 percent, followed by Grok 4.6 at 66.7 percent and GLM-5.3 at 62.2 percent.
Cost paints a different picture, and a pretty sharp one. An average index task costs $3.26 with GPT-6 Astra and $7.63 with Claude Fable 5.1. Lower down, GLM-5.3-Flash and GPT-5.6 Terra both score 42 points, but the Chinese model costs $0.25 per task versus $1.40 for OpenAI’s model. Artificial Analysis says four labs now sit on the Pareto frontier between intelligence and cost, with OpenAI holding most of that curve.
My take — AI-written commentary, not fact-checked reporting
Benchmark churn is doing a lot of work here. Rebuilding a leaderboard three times in a week is a clever way to keep vendors honest, but it also makes victory look a bit like weather: real, visible, and gone by lunch. OpenAI can enjoy the photo finish; everyone else gets to argue about the measuring tape.
Read more about this at: Trending Topics