TLDRocket
Sign in

Advancing the price-performance frontier with GPT-5.6

TLDR Dev

OpenAI just slashed GPT-5.6 pricing, with the small Luna model down 80% in cost.

OpenAI keeps chipping away at the cost of running its models, and the latest round of cuts on GPT-5.6 is the steepest yet. Luna, the smallest tier in the lineup, now costs $0.20 per million input tokens and $1.20 per million output tokens, an 80% drop from where it sat before. Terra, the mid-tier model, gets a more modest trim, down 20% to $2 per million input and $12 per million output.

These aren't isolated sticker changes. The lower rates carry over into Codex, OpenAI's coding assistant, and into the token quotas that ChatGPT Work customers draw down. So a business running a fleet of Codex agents or burning through Work seats gets the same discount without touching a config file.

The other change is arguably more interesting for anyone doing latency-sensitive work. Priority Processing, previously a flat premium for jumping the queue, is being replaced by something called Fast mode. On Sol, OpenAI's top-end model, Fast mode runs 2.5 times quicker than standard processing, but it costs twice as much. That's a very different trade than before: instead of paying extra just to skip the line, you're now paying for a genuinely faster inference path.

Taken together, this looks like OpenAI leaning harder into the price-performance argument it's been making against Anthropic and Google. Cheaper small models widen the net for high-volume, low-stakes tasks like classification or simple agent loops, while the Fast mode option gives enterprise customers a lever to pull when speed actually matters more than cost. It's a two-sided bet: undercut on the cheap end, upsell on the fast end.

My take

An 80% cut on Luna is the kind of number that gets headlines, but the real signal is Fast mode's 2x price for 2.5x speed — that's OpenAI finally admitting some customers don't care about cost per token, they care about seconds. Expect Anthropic and Google to answer with their own speed tiers within a quarter, because nobody wants to be the slow, cheap option in enterprise deals.

Read more about this at: TLDR Dev

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.