TLDRocket
Sign in

OpenAI cuts Luna API prices 80% and launches faster Sol inference

The Neuron Covered by 14 sources

OpenAI just slashed Luna's API price 80% and shipped a speedier Sol Fast mode. Translation: intelligence is getting cheaper and faster, on purpose.

OpenAI dropped the hammer on its own pricing this week. GPT-5.6 Luna's API rate fell 80%, down to $0.20 per million input tokens and $1.20 per million output tokens — numbers that make it one of the cheapest frontier-grade models on the market. Terra got a smaller haircut, 20%, landing at $2 and $12 respectively. And for anyone who wants speed over savings, there's now a Sol Fast mode that runs up to 2.5 times quicker for double the price. The same lower rates rolled out to ChatGPT Work and Codex on July 30, so this wasn't a quiet backend tweak — it hit every surface at once.

OpenAI says the cuts are self-funded, not charity. The company found production-kernel improvements and refined its speculative decoding, the trick where a model predicts several tokens ahead instead of generating one at a time. That lowered OpenAI's own serving costs enough that passing savings to developers apparently made business sense. Sam Altman, Greg Brockman, and Gavin Purcell all leaned on the same phrase in their commentary: price-per-intelligence. It's the metric OpenAI wants people using now, not raw capability scores.

The developer crowd noticed fast. OpenRouter's Shashank Goyal called Luna the most efficient dollar-per-token option currently available, and Cognition quickly updated its FrontierCode tooling to reflect the new math. Brockman went further, saying the lineup now sits right on the price-performance frontier — the sweet spot between cost and capability that most labs claim but few actually hit. Omar Sar's take was blunter: he called it intelligence too cheap to meter, then immediately flagged the catch — Luna can still act as a sub-agent through custom orchestration setups, but native Codex multi-agents v2 doesn't support that configuration yet. A separate thread from Tak found the same gap and floated a workaround using manual thread orchestration.

What makes the timing notable is how soon this came after launch. Weeks, not months. That kind of turnaround usually signals competitive pressure rather than routine optimization, and it's arriving right as Chinese open-weight models have climbed to roughly 48% of OpenRouter traffic, up from 20% a year ago, while U.S. models slipped to 32%. Cheaper compute, faster inference, aggressive repricing — it all reads like a company defending share, not just passing along efficiency gains out of generosity.

My take

Calling this "intelligence too cheap to meter" is cute marketing, but let's be honest about what's actually happening: OpenAI is defending its lunch money against a wave of cheap, capable open-weight models eating its market share overseas. I like cheaper tokens as much as the next developer, but the real story here isn't generosity, it's margin defense dressed up as a gift. If you want the actual price war, watch what open models do next, not what closed labs announce this week.

Read more about this at: The Neuron

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.