GPT-6 Astra Trails Top Models From Anthropic and Meta in Benchmarks
Trending Topics Jakob Steinschaden ● Covered by 2 sources
OpenAI launched GPT-6 Astra, but early tests don’t crown it. It ties GPT-5.6 on intelligence, while Anthropic and Meta still sit ahead.
Based on reporting by Trending Topics, Jakob Steinschaden — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
OpenAI has rolled out GPT-6 Astra as its new flagship model, and the launch came with the usual big words. Greg Brockman closed the briefing by invoking the “AGI era,” and even floated the idea that Astra might already count as artificial general intelligence. The first outside benchmark, though, is far less dramatic.
Artificial Analysis, the San Francisco-based benchmarking firm with an office in Melbourne, put Astra through its standard tests without any help from OpenAI. Its Intelligence Index covers math, science, programming, long-document reasoning and factual knowledge. On that measure, Astra scores 61 points — exactly the same as GPT-5.6 Sol — while Anthropic’s Claude Fable 5.1 is five points higher at 66, and Meta’s Muse Spark 1.3 also lands above OpenAI’s model.
The picture looks better in coding agents. In Artificial Analysis’ Coding Agent Index, Astra reaches 67 points in the Codex harness, roughly matching Claude Opus 5 and Fable 5 in Claude Code, and Muse Spark 1.3 in Muse Code. Fable 5.1 still leads that field with 70 points, but Astra’s efficiency changes the conversation. OpenAI has raised prices by 2.5x versus GPT-5.6 Sol, to $10 per million input tokens and $50 per million output tokens, yet Astra uses far fewer tokens in coding tasks — about one third of GPT-5.6 Sol’s count and about one fifth of Claude Opus 5’s.
That efficiency does not carry over everywhere. In the Intelligence Index, Astra saves around 10 percent in output tokens at max effort, which does not come close to offsetting the higher price. Artificial Analysis says the model ends up 75 percent more expensive per task than GPT-5.6 Sol there. In other words: a cleaner bill in coding, a nastier one for general intelligence.
The most striking gain shows up in AA-Omniscience, where the hallucination rate drops from 92 percent to 51 percent at max effort, while accuracy improves by four points. Some agent-style work also improves, with AA-Briefcase up about 80 Elo points. But not everything moves up. GDPval-AA v2 drops by roughly 80 Elo points, and Astra also slips by two to three points in τ³-Banking, SciCode and AA-LCR. The launch sounds like a turning point; the benchmark tables read more like a mixed bag with a sharper price tag.
My take — AI-written commentary, not fact-checked reporting
This is the part of AI release theatre that never gets old: announce the dawn of something huge, then let the tables do the rude talking. OpenAI seems determined to price confidence like luxury software, even when the model only clearly separates itself in some coding setups. The real story isn’t AGI cosplay; it’s that benchmark bragging rights are getting harder to buy on intelligence alone.
Read more about this at: Trending Topics