Grok 4.6 matched Fable 5 Max at an 85% discount. Downloadable models set that price.
The New Stack Matthew Burns
Grok 4.6, Qwen 3.8-Max and DeepSeek V4-Pro all landed in one day. The real story: downloadable models are crushing prices, so closed labs can’t charge much more.
Based on reporting by The New Stack, Matthew Burns — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Three frontier models showed up in about 24 hours, and the weird part wasn’t the speed. It was the pricing. Grok 4.6 launched on Wednesday, Qwen 3.8-Max arrived a few hours later, and DeepSeek V4-Pro followed on Thursday. All three sold themselves on cost, because the basic assumption now is that the capability race is already settled.
SpaceXAI framed Grok 4.6 as a “significant improvement over Grok 4.5 at the same price,” and the numbers back that up. Artificial Analysis put it at 61 on its Intelligence Index, five points above 4.5 and level with GPT-5.6 Sol. On GDPVal-AA, it reached 1,753 Elo, just ahead of Fable 5 Max at 1,741. Product analyst Aakash Gupta said the model still runs on the same 1.5 trillion parameters as 4.5, so the gains came from post-training rather than a bigger base model.
That matters because the API bill didn’t move. Grok 4.6 stays at $2 in and $6 out per million tokens. The model got better, but the serving cost stayed flat, which is why people keep talking about downloadable weights instead of benchmark glory. Two of this week’s three frontier launches came with weights you can download, and that keeps pushing the ceiling on what a closed lab can charge.
The comparison to Fable 5 Max makes the pressure obvious. Investor Gavin Baker ran the math and said Grok 4.6 is 80% cheaper on input tokens and 88% cheaper on output tokens than Anthropic’s model. He called it “Pareto dominant.” In other words, if a buyer can switch to a downloadable model, the premium for a closed one starts looking hard to justify.
That’s where the market seems to be going: not just cheaper models, but routing between models. Nvidia shipped an open router, NeMo Switchyard, alongside an open 30-billion-parameter model. Box CEO Aaron Levie’s point was simple: when the models get cheaper, value shifts to whoever decides which one does the job. The model is becoming plumbing. The chooser is becoming the product.
My take — AI-written commentary, not fact-checked reporting
This is the part the AI industry keeps pretending is temporary, and it isn’t. Once downloadable weights exist, “premium pricing” turns into a confidence trick with nicer typography. The real business is moving up a layer to routing, control and distribution, which is inconvenient for anyone who thought the model itself was the moat.
Read more about this at: The New Stack