TLDRocket
Sign in

Price per 1M tokens is meaningless

janilowski.pl Covered by 2 sources

Comparing AI models by price per token is basically meaningless. A pricier model can finish the same task for half the cost of a cheaper-looking one.

Based on reporting by janilowski.pl — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Everyone loves quoting a model's price per million tokens like it's a fuel-efficiency rating, but the number barely tells you anything. Each frontier lab runs its own tokenizer, so the same chunk of text gets chopped up differently depending on whose model you're feeding it to. A passage that comes out to 160 tokens for GPT-4o balloons to 200 tokens for GPT-4's 1106-preview version, and that's within a single company's own lineup. Anthropic recently tweaked its tokenizer too, and the same text now splits into 30 percent more tokens than before, which quietly functions like a price increase even though the sticker price didn't move.

Then there's the bigger problem hiding underneath: token efficiency. A lot of AI usage today involves models doing extended, often hidden reasoning before they spit out an answer, and that thinking gets billed at the same rate as the visible output. Two models can produce equally good answers while one of them burns through wildly more tokens getting there, and price per token has no way of capturing that difference.

Looking at cost per completed task on the Artificial Analysis benchmark makes the gap obvious. GPT-5.5 xhigh is nominally more expensive per token than Claude Opus 4.8 max, yet it finishes benchmark tasks for roughly half the cost. GLM-5.2 max looks like a bargain on paper, priced several times cheaper per token than both GPT and Claude offerings, but its cost per task doesn't shrink nearly as much, which points to it being less token-efficient than the Western frontier models it's often pitched against. DeepSeek V4 Pro max stands out as the real outlier here, scoring lower on raw intelligence but landing an extremely low cost per task, making it look like the most efficient option in the bunch. Claude Sonnet 5 max is the strange one: it scores worse than Opus 4.8 while costing more per task, thanks to poor token efficiency, and it's genuinely unclear what problem it's meant to solve. Fable 5, meanwhile, charges more than three times what GPT-5.5 does for what looks like only a modest capability bump.

None of this shows up if you're just eyeballing dollars per million tokens. Skipping the cost-per-task math means picking models based on a number that doesn't reflect what you're actually paying to get work done, and ending up with worse performance for more money.

My take — AI-written commentary, not fact-checked reporting

Sticker prices on AI models are basically marketing copy at this point, and anyone still shopping by dollars-per-million-tokens deserves the inflated bill they'll eventually get. DeepSeek's showing here, cheap per token and cheap per task, is the kind of result that should worry the labs charging a premium for marginally better benchmark scores. And Claude Sonnet 5 costing more per task than its own sibling Opus 4.8, while scoring worse, smells less like an engineering choice and more like a pricing trick nobody's bothered to explain.

Read more about this at: janilowski.pl

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.