TLDRocket
Sign in

Tokenomics: Why making AI pay is tricky

BBC Covered by 2 sources

AI companies and their customers can't agree on how to price AI agents, because nobody can predict how many tokens the things will actually burn through. That unpredictability is turning simple software billing into a guessing game with real money on the line.

Based on reporting by BBC — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Ask ChatGPT to plan your holiday and you're getting a genuinely good deal, subsidized by hundreds of billions of dollars that Microsoft, Google and Anthropic have poured into building these models. But somebody has to pay that money back eventually, and that's where things get messy. The bill comes down to tokens, the small chunks of text an LLM breaks a request into before spitting out an answer. The trouble is that the same question can produce wildly different numbers of tokens depending on phrasing, the model used, or just the roll of the dice, since these systems aren't deterministic in the way a calculator is.

That unpredictability turns into a real financial headache once businesses start stacking multiple AI agents together to handle tasks automatically. Goldman Sachs reckons token consumption by businesses will jump 24-fold between 2026 and 2030, hitting 120 quadrillion tokens a month. Individual tokens have gotten cheaper, sure, but volume is exploding fast enough to erase those savings. Uber reportedly blew through an entire year's coding-token budget in just a few months this year. Microsoft, of all companies, has had to pull back its own engineers' use of third-party coding tools.

Simon Gooch of Saviynt put it bluntly: locking anyone into a fixed price for the next one, two or three years makes no sense right now, because nobody actually knows what usage will look like by then. Some smaller firms are dodging the problem entirely by running flat-fee personal accounts instead of enterprise contracts, essentially flying under the radar, according to smartR AI founder Oliver King-Smith. He doesn't think that lasts. Once the big AI vendors face shareholder pressure to turn a profit, he expects a crackdown on exactly that kind of workaround.

And the math gets weirder the deeper you go. Companies budgeting for AI often only plan for the token costs tied directly to writing code, forgetting the extra spend needed for testing, security checks or safety guardrails, says LSE's Will Venters. Adding another AI agent takes one click; adding another human employee takes a hiring process. That asymmetry makes it dangerously easy for token costs to creep up unnoticed until a company gets its bill or simply runs dry.

Down the chain, firms selling AI-powered products are stuck figuring out how to pass these shifting costs to their own customers. Sumo Logic's Bill Peterson admits nobody has really cracked it yet, joking that internal pricing conversations have gotten pretty entertaining. Flat price hikes, pay-per-result models, bundled incident pricing, all are on the table. But whatever pricing scheme a company lands on could be wrecked overnight if the big LLM providers shift their own rates again, something Peterson says happens every couple of months. Customers hate that kind of instability. Nobody builds a budget around a number that keeps moving.

My take — AI-written commentary, not fact-checked reporting

This is what happens when an entire industry builds its business model on top of infrastructure it doesn't control and can't predict, and it's a preview of pain to come for anyone selling AI agents rather than just AI chat. The vendors love unpredictable usage because it means unpredictable, growing bills; customers hate it for exactly the same reason. Expect the smart money to move toward flat-rate, capped-usage products sooner than the big platforms would like, because businesses will not tolerate a line item that swings 20x without warning.

Read more about this at: BBC

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.