TLDRocket
Sign in

GPT-6 Astra Now on Par With Claude Fable 5.1 in Updated Artificial Analysis Index

Trending Topics Jakob Steinschaden Covered by 14 sources

GPT-6 Astra just tied Claude Fable 5.1 at the top of Artificial Analysis’ new index. That jump came after the benchmark was rewritten twice in one week.

Based on reporting by Trending Topics, Jakob Steinschaden — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Artificial Analysis has bumped its Intelligence Index to version 4.3, and the top slot is now a dead heat. GPT-6 Astra (max) from OpenAI and Claude Fable 5.1 (max with fallback) from Anthropic both land on 53 points. Right behind them sit Claude Opus 5 on 51, Claude Fable 5 on 50, Meta’s Muse Spark 1.3 on 48, and GPT-5.6 Sol on 47.

The bigger story is how quickly OpenAI’s model moved. Within a few days, and across two index updates, GPT-6 Astra went from fifth place to first. Less than a week separates its debut from this result, and in that short span Artificial Analysis changed the test mix twice.

Two benchmarks were swapped or upgraded. Terminal-Bench jumps from 2.1 to 4.0 and uses 66 tasks to check whether an agent can handle complex work through the terminal, across software, machine learning, science, operations, security, hardware and media. Astra leads that one at 59.1 percent, ahead of Claude Fable 5.1 at 52.0 percent, Claude Opus 5 at 49.0 percent and GPT-5.6 Sol at 39.9 percent.

The new AutomationBench-AA replaces the old banking test 𝜏³. Built by Zapier, it spans finance, HR, marketing, operations, sales and support, with 657 tasks set inside simulated business apps. Agents have to figure out the right APIs on their own, and any guardrail breach zeros out that task. Astra again tops the chart at 68.5 percent, followed by Grok 4.6 at 66.7 percent and GLM-5.3 at 62.2 percent.

Cost paints a different picture, and a pretty sharp one. An average index task costs $3.26 with GPT-6 Astra and $7.63 with Claude Fable 5.1. Lower down, GLM-5.3-Flash and GPT-5.6 Terra both score 42 points, but the Chinese model costs $0.25 per task versus $1.40 for OpenAI’s model. Artificial Analysis says four labs now sit on the Pareto frontier between intelligence and cost, with OpenAI holding most of that curve.

My take — AI-written commentary, not fact-checked reporting

Benchmark churn is doing a lot of work here. Rebuilding a leaderboard three times in a week is a clever way to keep vendors honest, but it also makes victory look a bit like weather: real, visible, and gone by lunch. OpenAI can enjoy the photo finish; everyone else gets to argue about the measuring tape.

Read more about this at: Trending Topics

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.