TLDRocket
Sign in

Open-weight models now handle a majority of tokens on Vercel’s AI Gateway. But Anthropic still takes 64% of the spend.

The New Stack Paul Sawers Covered by 2 sources

Open-weight models now do 56% of tokens on Vercel’s AI Gateway. But Anthropic still takes 64% of the spend.

Based on reporting by The New Stack, Paul Sawers — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Open-weight models have crossed a line that matters: they’re no longer a niche share of production traffic. On Vercel’s AI Gateway, they handled 56% of all tokens routed through the service in August, the first month they’ve held a majority. That share was only 7% in December 2025, climbed to 13% by April, and then kept rising month after month.

Vercel’s gateway is a useful place to watch this shift. The company says the service routes tens of trillions of tokens each month, sitting between apps and model providers while tracking both usage and cost. That makes token volume a pretty good stand-in for how much inference work is actually being sent to different models, even if it says nothing by itself about dollars.

And the dollars tell a different story. Open-weight models accounted for just 14 cents of every estimated dollar spent through AI Gateway in August. Anthropic alone took 64 cents of every dollar, and Vercel says its share has never dropped below 61% in any month since December 2025. Meanwhile, the average price per token across the gateway fell 23.2% in August, the third straight monthly decline.

There’s movement inside that Anthropic bucket, though. Fable 5 fell from 13.2% of total gateway spend in July to 4.9% in August, while Opus 5 rose to 22.5%. Vercel says 90% of teams using Fable reduced usage, and more of that work moved to Opus 5 than to any other model. The newer model did similar jobs at roughly half the price, which is a neat reminder that customers often stay loyal to the lab, not the name on the box.

That same pattern showed up with other providers too. Z.ai’s GLM-5.3-Flash reached three times the daily volume of GLM-5.2 within five days of launch, while more than three-quarters of the volume lost by Google’s Gemini 3 Flash moved to other providers, including OpenAI and Anthropic. The market is getting less sticky by the month. Models can win or lose fast. The lab only keeps the money if the replacement is good enough.

My take — AI-written commentary, not fact-checked reporting

This is the part of the AI business that gets missed when everyone is busy chanting “open vs closed” like it’s a football derby. Users don’t seem to care about ideology; they care about price, fit, and whether the model actually does the job. Brand loyalty is thin. Model loyalty is the real moat, and it’s a much uglier thing to defend.

Read more about this at: The New Stack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.