TLDRocket
Sign in

GPT-6 Sol vs. Claude Opus 5.5: Is cheaper important when results aren’t consistent?

The New Stack Jessica Wachtel ● Covered by 8 sources

OpenAI says GPT-6 Sol beats Claude Opus 5 on workflow tasks for far less. In repeat tests, Sol was cheaper and faster, but Opus 5.5 was steadier.

Based on reporting by The New Stack, Jessica Wachtel — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

OpenAI launched GPT-6 Sol on September 22 and pitched it as a cheaper model that outperforms Claude Opus 5 on business workflow tasks. Its own marketing puts Sol at 33.2% on Zapier’s AutomationBench, versus 26.9% for Opus 5, at roughly 9% of the cost per task. But Anthropic released Opus 5.5 the same day, so the comparison got awkward fast.

That awkwardness is the whole story here. The first pass was a wash: both models handled the usual developer chores about equally well, so the test changed. Instead of asking which one could solve the prompt once, the focus became whether it could do it again and again. Each model was run five times through identical API calls at maximum effort.

The three tests were not toy problems. One had the models sort 40 failed CI jobs using a runbook. Another handed them 3,664 lines of outage logs from five services and asked seven postmortem questions. The third asked for a dependency resolver from a two-page spec, with a hidden suite of 120 tests waiting underneath. Both models were perfect on the first round of harder versions, so the repeat runs became the tiebreaker.

On CI triage, both models went 5 for 5. Sol was much faster and far cheaper, averaging 18 seconds and $0.02 per run, while Opus 5.5 averaged 1 minute 27 seconds and $0.24. That lined up closely with OpenAI’s price claim, even though Opus 5.5 is cheaper per token than Opus 5, because it used far more output tokens on the test.

The other two tasks tell a less flattering story for Sol. On the incident logs, Opus 5.5 stayed perfect across all five runs. Sol did not: it nailed two runs, then missed one customer in two others and counted 27 failed checkouts instead of 28 in another. On the resolver spec, Opus 5.5 again ran the table, while Sol passed four times and once shipped a stray closing parenthesis that broke the module on import. Across all 15 runs, Sol was faster and cheaper every time, but it was only perfect 12 times. Opus 5.5 was perfect 15 out of 15.

My take — AI-written commentary, not fact-checked reporting

Cheap only matters if the output survives contact with reality. Sol looks great for high-volume work with a second set of eyes; for anything where a missed detail turns into real pain, the steadier model wins. Price is nice, but consistency is the feature.

Read more about this at: The New Stack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.