Advancing the price-performance frontier with GPT-5.6
TLDR Dev
OpenAI just slashed GPT-5.6 pricing, with the small Luna model down 80% in cost.
OpenAI keeps chipping away at the cost of running its models, and the latest round of cuts on GPT-5.6 is the steepest yet. Luna, the smallest tier in the lineup, now costs $0.20 per million input tokens and $1.20 per million output tokens, an 80% drop from where it sat before. Terra, the mid-tier model, gets a more modest trim, down 20% to $2 per million input and $12 per million output.
These aren't isolated sticker changes. The lower rates carry over into Codex, OpenAI's coding assistant, and into the token quotas that ChatGPT Work customers draw down. So a business running a fleet of Codex agents or burning through Work seats gets the same discount without touching a config file.
The other change is arguably more interesting for anyone doing latency-sensitive work. Priority Processing, previously a flat premium for jumping the queue, is being replaced by something called Fast mode. On Sol, OpenAI's top-end model, Fast mode runs 2.5 times quicker than standard processing, but it costs twice as much. That's a very different trade than before: instead of paying extra just to skip the line, you're now paying for a genuinely faster inference path.
Taken together, this looks like OpenAI leaning harder into the price-performance argument it's been making against Anthropic and Google. Cheaper small models widen the net for high-volume, low-stakes tasks like classification or simple agent loops, while the Fast mode option gives enterprise customers a lever to pull when speed actually matters more than cost. It's a two-sided bet: undercut on the cheap end, upsell on the fast end.
My take
An 80% cut on Luna is the kind of number that gets headlines, but the real signal is Fast mode's 2x price for 2.5x speed — that's OpenAI finally admitting some customers don't care about cost per token, they care about seconds. Expect Anthropic and Google to answer with their own speed tiers within a quarter, because nobody wants to be the slow, cheap option in enterprise deals.
Read more about this at: TLDR Dev
Related stories
Advancing the price-performance frontier with GPT‑5.6
Simon Willison · 3 days ago ·
16
Advancing the price-performance frontier with GPT-5.6
OpenAI Blog · 4 days ago ·
13
[AINews] GPT 5.6 price cut by 20%-80%: Cost of GPT 5.4 Intelligence dropped 13x in 4 months due to GPT 5.6 recursive self-optimization
Latent Space · 3 days ago ·
1