GPT-6 Astra aced the hardest AI benchmark. The asterisk matters more than the score.
The New Stack Amanda Caswell ● Covered by 2 sources
OpenAI says GPT-6 Astra scored 98.6% on ARC-AGI-3. The catch: it was tested in a setup that may have helped, so the asterisk matters almost as much as the score.
Based on reporting by The New Stack, Amanda Caswell — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
OpenAI has a new bragging point: GPT-6 Astra scored 98.6% on ARC-AGI-3, a benchmark built to trap models in unfamiliar interactive settings. Six months earlier, frontier models were barely clearing 1% while humans could work through the tasks. Astra’s predecessor, GPT-5.6 Sol, was at 7.8% in OpenAI’s numbers. That is not a small improvement. It is the kind of jump that makes the old score look like a typo.
But ARC-AGI-3 is designed to test whether a model can figure out the rules of a new environment, not just regurgitate patterns from training data. So the setup matters. Astra was evaluated through OpenAI’s Responses API harness, with two settings adjusted to better match how the model behaves in real use. OpenAI says those tweaks were not made for ARC-AGI-3 specifically, though the company also says the other models in the comparison were tested in different setups. On a benchmark like this, that is not a footnote. It is part of the result.
The headline score also doesn’t stand alone. OpenAI says Astra hit 97.6% on FrontierMath Tier 4, 100% on ExploitBench, and 99.2% on SRE-Bench with four attempts. Terminal-Bench Science showed a much bigger leap, from 22.4% for Sol to 64.6% for Astra. The company warns against rolling all of that into one neat measure of intelligence, which is probably wise. Still, the spread says Astra is taking on a wider range of tasks than the model it replaced.
That shows up in the demos, too. OpenAI says Astra works directly inside tools such as KiCad, Power BI, and Unity. An experimental Codex feature gives it notes and search across earlier context when a job outlasts a single context window. On offline OSWorld 2.0, Astra scored 72.6% and took about 40 minutes per task, versus Sol’s 65.7% and roughly 75 minutes. Speed matters here. So does endurance.
The most interesting claims are the math ones. OpenAI says Astra helped with two new findings about gaps between prime numbers. One bound, which mathematician Julia Stadlmann had already pushed from 246 to 240, was lowered again to 186 with Astra involved. In another case, the model helped improve part of a bound that hadn’t moved in more than 80 years. But OpenAI doesn’t say exactly what Astra found on its own, what the researchers supplied, or how the work was divided up. That leaves the result intriguing without making it proof of anything grander.
And that is the real story here. Astra looks like a model that can handle longer, messier, more useful work than OpenAI’s last one, and it even avoids some of the worst behavior in internal tests. But the benchmark, the harness, and the human scaffolding still matter. If this is what progress looks like now, it is less a clean leap to AGI than a messy, very expensive argument about what counts as intelligence in the first place.
My take — AI-written commentary, not fact-checked reporting
The industry keeps trying to turn model scores into destiny, and then acts surprised when the setup turns out to be half the story. Astra looks genuinely stronger, but the benchmark theater is getting old fast. The useful pattern is simpler: systems are improving, oversight still matters, and nobody should mistake a very good score for an argument that the robots have become wise.
Read more about this at: The New Stack