TLDRocket
Sign in

The Sequence Opinion #909: Return on Token: The New Economics of AI-Native Engineering

Substack Jesus Rodriguez

Opinion — commentary, not a factual news event.

Companies are now tracking AI coding agents by how many tokens they burn, treating usage like a productivity metric. Turns out, more tokens spent doesn't mean better engineering — it just means bigger bills and a new vanity metric to game.

There used to be a simple formula for engineering output: people, multiplied by their time, multiplied by their skill. You could staff a project, forecast a roadmap, and mostly trust that headcount mapped to capacity. That math is breaking down fast, and the culprit is the swarm of AI agents now doing real engineering work inside companies.

Today an engineer can spin up four or five agents at once — one hunting a production bug, another writing test coverage, a third sketching a new architecture, a fourth drafting docs — and let them run in parallel for hours. None of these agents show up in a headcount spreadsheet. None of them negotiate equity or sit through a planning offsite. But they do burn tokens, and lots of them, which means organizations now have a second workforce that scales elastically with spend rather than with hiring.

Naturally, this has created a new obsession: token maxing. Because tokens are countable and countable things get dashboards, teams have started treating token consumption the way an earlier generation treated lines of code — as a proxy for productivity, even when it isn't one. The comparison to lines-of-code metrics isn't flattering. That earlier era taught the industry that optimizing for a raw output number tends to produce bloated, low-quality results rather than genuinely better software.

The deeper problem is that nobody has settled on what a 'good' token-to-output ratio actually looks like. Is heavy token usage a sign of a thorough, well-instrumented agent doing rigorous work, or a sign of an inefficient one looping and re-checking itself into a bigger bill? Right now, most organizations can't tell the difference, and that ambiguity is exactly why the metric is spreading — it's easy to measure and hard to interpret, which is often how bad KPIs are born.

What's really happening is a shift in how engineering capacity gets planned and budgeted. Compute spend is becoming a line item as strategically important as hiring plans once were, and companies that figure out how to measure agent output by actual value delivered — bugs fixed, features shipped, incidents resolved — rather than tokens burned, are going to have a real advantage over everyone else still staring at a token dashboard congratulating themselves for spending more.

My take

Token maxing is what happens when an industry mistakes a byproduct for a benchmark, and anyone who's watched a KPI get gamed before knows exactly how this movie ends: teams will optimize for the number, not the outcome, and someone's quarterly review will suffer for it. The real signal was never going to be how many tokens an agent burns — it's whether the bug actually got fixed. Companies chasing token dashboards instead of value delivered are just rebuilding the lines-of-code fiasco with a shinier unit of measurement.

Read more about this at: Substack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.