TLDRocket
Sign in

After blowing its entire 2026 AI budget in months, Uber CTO says ‘We’re coming to the end of the so-called 'tokenmaxxing' era’

Fortune Sasha Rogelberg Covered by 4 sources

Uber blew its entire 2026 AI budget in just a few months after pushing employees to max out Claude Code usage, even ranking them on leaderboards. Now its CTO says the 'tokenmaxxing' era is over, having quadrupled AI adoption while actually cutting cost per token through caching and smarter defaults.

Based on reporting by Fortune, Sasha Rogelberg — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Uber's CTO Praveen Neppalli Naga had a bit of a reckoning earlier this year. After telling staff to go wild with Anthropic's Claude Code, complete with leaderboards ranking engineers by usage, the company torched its entire 2026 AI budget in a matter of months. That's the kind of number that gets a CFO's attention fast.

So Naga went back to the drawing board, and what he came back with is a case study in the difference between throwing money at a problem and actually engineering your way out of it. Uber didn't cut access. Instead it quadrupled the number of employees using frontier AI tools while driving down the cost per token, largely by improving prompt caching, tweaking default model settings, evaluating newer models for efficiency, and giving engineers visibility into their own hourly AI spend. Costs should have climbed as usage exploded. They didn't. Naga frames this as the end of what he's calling the tokenmaxxing era, where volume of usage was the metric that mattered.

The backdrop here is not flattering for the AI-productivity narrative. Deutsche Bank's Jim Reid recently cautioned that broad productivity gains from AI remain years off, and the numbers back him up: Magnificent Seven profit margins jumped from 15% to 25% between early 2023 and early 2026, while the rest of the S&P 500 crawled up just 10% over the same stretch. Even inside Uber, President Andrew Macdonald admitted in May that the company still can't draw a clean line between AI spending and shipping measurably better features for riders.

And there's a nastier problem lurking underneath all this efficiency talk: Jevons paradox. Cheaper tokens don't necessarily mean less spending, they tend to mean more agents, more automated workflows, more generated code, until aggregate spend rises anyway. The Silicon Data Token Expenditure Index shows token prices down more than 90% since 2023, yet LLM spending has doubled since late last year. Bain found something similar, with per-token costs halving from December 2024 to 2025 while total tokens consumed grew 450%. Naga didn't say whether Uber's overall compute usage went up or down, only that the company is now optimizing for quality over sheer volume. Whether that holds once the leaderboards are gone is the real test.

My take — AI-written commentary, not fact-checked reporting

Calling this the end of tokenmaxxing is a bit premature when Uber won't even say if total compute spend went up or down, just that it's more efficient per unit. That's the classic Jevons paradox dodge every AI vendor and enterprise is currently performing: efficiency gains get announced loudly while aggregate spending quietly climbs, because nobody's actually incentivized to use less, just to look smarter while using more.

Read more about this at: Fortune

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.