Agentic AI is breaking the token meter, and enterprises need a plan for what comes next
SiliconANGLE Zeus Kerravala ● Covered by 4 sources
Agentic AI can burn 10 to 100 times more tokens per task than a simple chat call. That makes per-token pricing look cheap right up until production starts.
Based on reporting by SiliconANGLE, Zeus Kerravala — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Per-token pricing was supposed to make AI easy to try. It did. But a new Futurum report argues that the same billing model becomes a trap once agents move from demos to real work.
The report, “The Off Ramp From Per-Token Pricing,” says agentic AI can consume 10 to 100 times more tokens per task than a plain inference call. That matters because the expensive part of AI is no longer the pilot. It’s the thing people actually want to keep using.
Zeus Kerravala, who wrote the piece for SiliconANGLE, says the pattern keeps showing up with CIOs and CFOs: the most successful AI projects are often the ones that blow past their budgets. He points to one organization that planned to spend $1 million for a year and burned through it in three months. That’s not a model problem. That’s success meeting a meter that never learned the difference between a test and a production system.
Futurum’s numbers make the pressure clearer. The firm expects agent and reasoning inference to grow by 219% this year, and total inference spending to rise from $120 billion in 2025 to $885 billion by 2030. Against that backdrop, a pricing model that scales directly with usage starts to look less like convenience and more like a tax on useful software.
The report also shows that enterprises have already been voting with their infrastructure. In a survey of 824 AI decision-makers, reserved and owned infrastructure made up 66% of AI compute consumption, versus 19% for on-demand cloud. Another 59% said they mainly run AI workloads outside hyperscaler public clouds, in their own data centers, colocation facilities, or with bare-metal providers.
That shift is not really a rejection of the cloud. It’s a sign that companies are getting smarter about where steady workloads belong. Amberd.ai’s setup on QumulusAI bare metal is the clearest example in the report: an eight-GPU Nvidia H200 server split into four virtual environments, with customers tiered by latency tolerance. The economics only work because the hardware stays busy.
And that’s the real point. Reserved infrastructure can beat variable pricing, but only when teams know their workloads well enough to keep utilization high. Futurum recommends reserved bare metal for sustained workloads with predictable utilization above roughly 60%, while also warning that this approach needs more custom engineering. Many enterprises simply don’t have that skill set yet.
My take — AI-written commentary, not fact-checked reporting
The industry loves to sell AI like electricity and bill it like a taxi. That works fine for experiments and ruins the budget once agents start doing real work. The grown-up move is boring: measure cost per task, not cost per token, and stop pretending the meter is neutral.
Read more about this at: SiliconANGLE