TLDRocket
Sign in

OpenAI quietly boosts some of Astra's evaluation metrics, and continues to change others post-launch

Fortune Emily Forlini Covered by 10 sources

OpenAI kept changing Astra’s benchmark scores after posting its launch blog. Some of the edits made Astra look better, while rivals’ numbers moved around too.

Based on reporting by Fortune, Emily Forlini — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

OpenAI’s GPT-6 Astra launch turned into a messy little lesson in how slippery AI benchmarks can be. The company first planned to publish the blog at 2 p.m. ET on Sept. 3, but the post kept disappearing, reappearing, and breaking for readers for almost two more hours. By the time it was stable, some of the evaluation numbers had already changed — and a few kept changing after that.

The most glaring shift was Astra’s hallucination rate. In the first archived version of the post, it sat at 4.2%. Later it was cut to 2%, while GPT-5.6 Sol dropped from 12.2% to 9.4%. Now those figures are back to 4.2% and 12.2%. OpenAI also altered an internal cybersecurity score for Sol, moving it from 5.5% to 11.5%, and said it is looking into reverting that because the higher number reflected a reasoning level not commercially available for the model.

The math scores moved in a way that briefly made Astra look even stronger. Astra itself stayed at 97.6% on FrontierMath Tier 4 (v2), but Anthropic’s Fable 5.1 slipped from 87.8% to 78% before rising again to 83%, while Sol moved from 83% to 80.5% and back to 83%. On ARC-AGI-3, Astra’s score was 98.6% in a draft sent to reporters and 99.99% in the live blog. OpenAI also pointed out that the Arc Prize Foundation measured Astra at 99.9% with a particularly powerful harness, but only 63% with the benchmark’s standard harness.

And the changes weren’t all flattering to Astra. Its coding score in the blog crept from 57.7% to 57.9%, while Anthropic models improved on HealthBench Professional. Claude Fable 5.1 went from 56.6% to 58.1%, and Opus 5 from 54.5% to 56.4%. OpenAI says evaluation scores are the “maximum at any effort,” but the whole episode shows why benchmark numbers are useful, messy, and easy to argue over all at once.

My take — AI-written commentary, not fact-checked reporting

This is exactly why benchmark theater keeps getting worse: every company says it’s just adjusting for “best estimate,” and every launch starts to look like a spreadsheet with stage makeup. OpenAI isn’t alone here, but when the numbers keep moving after publication, the public is right to assume the scoreboard is part lab report, part sales pitch. Benchmarks still matter; they just shouldn’t be treated like scripture carved on a server wall.

Read more about this at: Fortune

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.