TLDRocket
Sign in

Chinese AI competitors may have forced OpenAI’s hand on pricing

The New Stack Amanda Caswell Covered by 14 sources

OpenAI just slashed API prices on two GPT-5.6 models, only three weeks after launch. Luna's down 80%, and it's likely Chinese rivals like Moonshot forced the move.

Based on reporting by The New Stack, Amanda Caswell — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Three weeks. That's how long OpenAI let GPT-5.6's prices sit before cutting them. On Thursday, Sam Altman announced that GPT-5.6 Luna now costs $0.20 per million input tokens and $1.20 per million output tokens, down from $1 and $6 — an 80% drop. Terra got a smaller haircut, falling from $2.50/$15 to $2/$12 per million tokens. Sol, the flagship reasoning model, kept its price at $5/$30, though it did pick up a new Fast mode that runs 2.5x quicker for double the cost.

OpenAI frames this as an infrastructure win. A day before the price cut, the company published an engineering rundown detailing rewritten GPU kernels that trimmed serving costs by roughly 20%, a redesigned speculative decoding system for Sol that boosted token generation efficiency by more than 15%, and heavier use of prompt caching in its agent runtime to avoid recomputing the same context over and over. All real gains, and all things that take engineering time — which is exactly why the timing here feels off. Vendors typically sit on a new model family's pricing for months, not weeks.

The more interesting story is what's happening outside OpenAI's walls. Open-weight models out of China, from labs like Moonshot, have gotten good enough that they're no longer just cheap alternatives — they're viable production tools. When a company is burning through billions of tokens a day running agentic workflows that involve dozens or hundreds of model calls per task, the difference between a benchmark-topping model and a merely competent one stops mattering as much as the bill at the end of the month. Suddenly the case for self-hosting an open model, or routing simple tasks to a cheaper API, gets a lot stronger.

That's the pressure OpenAI is responding to, whether it says so directly or not. Luna's 80% cut looks less like generosity and more like an attempt to close a gap that Chinese labs pried open by packing solid capability into leaner, cheaper models. And OpenAI isn't alone in scrambling — Anthropic has been rolling out its own pricing tweaks and premium tiers as enterprise customers push bigger agentic workloads into production, and OpenAI itself just raised usage limits on Sol for ChatGPT Work and Codex after coding sessions burned through allowances faster than planned.

What this points to is a shift in how these companies compete. Model quality used to be the whole game. Now serving efficiency is becoming just as important, because every percentage point shaved off inference cost can turn into a headline price cut that keeps developers from bothering with alternatives. Engineering that used to be an internal cost center is now a marketing lever.

My take — AI-written commentary, not fact-checked reporting

This is China's open-weight push working exactly as intended — not by winning benchmark charts, but by making Western labs sweat over unit economics. I'd bet we see more surprise mid-cycle price cuts from OpenAI and Anthropic before the year's out, and I don't think that's a bad thing for anyone except their margins.

Read more about this at: The New Stack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.