Artificial Analysis benchmarks GPT-6 Astra vs other agent models
Artificial Analysis ● Covered by 5 sources
GPT-6 Astra beats or matches rival coding models, but it costs a lot more. It’s sharper on tokens and hallucinations, yet the price jump blunts the win.
Based on reporting by Artificial Analysis — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Artificial Analysis has put GPT-6 Astra through its Coding Agent and Intelligence indices, and the model comes out looking split-screen. On coding, it looks genuinely strong: Astra matches Fable 5 in the Coding Agent Index while using far fewer tokens, and at max effort it lands about the same price as GPT-5.6 Sol while scoring two points higher. On intelligence work, the story is messier. It matches GPT-5.6 Sol on the index score, but the newer model is also much pricier to run.
The biggest number in the report is the price change. GPT-6 Astra is now listed at $10 per million input tokens and $50 per million output tokens, up from $4 and $20. That is a 2.5x jump across the board. The model still keeps the same cache-read discount and cache-write premium, but those pricing details do not soften the main point: the new model is more expensive, even when it is spending fewer tokens.
Artificial Analysis says Astra uses about one third of the tokens GPT-5.6 Sol needs in the Codex harness at max effort, and about one fifth of what Claude Opus 5 uses there. That token thrift matters because it pushes Astra onto the cost-efficiency frontier in coding. In plain English, it is doing more with less, and that helps it land beside models like Claude Opus 5 and Fable 5 in the Coding Agent Index, where Fable 5.1 still leads with a score of 70.
The intelligence benchmark tells a different tale. Astra matches GPT-5.6 Sol at 61, but trails Fable 5.1 and Meta’s Muse Spark 1.3. It does shave about 10% off output tokens at max effort, yet the price increase still makes it 75% more expensive per task than its predecessor. So the efficiency story is real, but not enough to wash out the bill.
There are real quality gains tucked in there, too. Artificial Analysis says Astra cuts hallucination rate in its AA-Omniscience test from 92% to 51% at max effort, while also lifting accuracy by 4 points. It also gains about 80 points in AA-Briefcase, a long-horizon knowledge-work test built around multi-week projects, many linked tasks, and thousands of source files. But the model is not climbing in a straight line: Presentation Quality Elo falls, GDPval-AA v2 drops by about 80 Elo points, and a few other evaluations also move the wrong way.
So Astra looks like a model that is better at some hard agent tasks, better at staying on task without wandering into hallucination, and worse at pretending price doesn’t exist. That last part is usually where the marketing gets quiet.
My take — AI-written commentary, not fact-checked reporting
This is the usual AI industry trick: sell efficiency, then invoice for ambition. GPT-6 Astra looks like a real upgrade for coding, but the 2.5x price hike means the model is being pushed as progress while the spreadsheet does the heavy lifting. Open models keep getting told to catch up; closed models keep showing why buyers should read the bill first.
Read more about this at: Artificial Analysis