TLDRocket
Sign in

Advancing the price-performance frontier with GPT-5.6

OpenAI Covered by 2 sources

OpenAI just slashed GPT-5.6 pricing, with the small Luna model down 80% in cost.

Based on reporting by OpenAI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

OpenAI keeps chipping away at the cost of running its models, and the latest round of cuts on GPT-5.6 is the steepest yet. Luna, the smallest tier in the lineup, now costs $0.20 per million input tokens and $1.20 per million output tokens, an 80% drop from where it sat before. Terra, the mid-tier model, gets a more modest trim, down 20% to $2 per million input and $12 per million output.

These aren't isolated sticker changes. The lower rates carry over into Codex, OpenAI's coding assistant, and into the token quotas that ChatGPT Work customers draw down. So a business running a fleet of Codex agents or burning through Work seats gets the same discount without touching a config file.

The other change is arguably more interesting for anyone doing latency-sensitive work. Priority Processing, previously a flat premium for jumping the queue, is being replaced by something called Fast mode. On Sol, OpenAI's top-end model, Fast mode runs 2.5 times quicker than standard processing, but it costs twice as much. That's a very different trade than before: instead of paying extra just to skip the line, you're now paying for a genuinely faster inference path.

Taken together, this looks like OpenAI leaning harder into the price-performance argument it's been making against Anthropic and Google. Cheaper small models widen the net for high-volume, low-stakes tasks like classification or simple agent loops, while the Fast mode option gives enterprise customers a lever to pull when speed actually matters more than cost. It's a two-sided bet: undercut on the cheap end, upsell on the fast end.

My take — AI-written commentary, not fact-checked reporting

An 80% cut on Luna is the kind of number that gets headlines, but the real signal is Fast mode's 2x price for 2.5x speed — that's OpenAI finally admitting some customers don't care about cost per token, they care about seconds. Expect Anthropic and Google to answer with their own speed tiers within a quarter, because nobody wants to be the slow, cheap option in enterprise deals.

Read more about this at: OpenAI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.