TLDRocket
Sign in

Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents

NVIDIA Shruti Koparkar Covered by 2 sources

NVIDIA says Vera Rubin NVL72 can do up to 30x more agent work per watt than GB300 NVL72. That matters because agentic AI burns far more tokens than chat, so power bills become the bottleneck.

Based on reporting by NVIDIA, Shruti Koparkar — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

NVIDIA is leaning hard into a simple argument: agentic AI is not chat with extra steps, and the hardware should be judged that way too. The company cites OpenRouter data showing agentic workloads use 15x more tokens than a plain chat request, then points to its own measurements on real coding sessions to argue that the next bottleneck is power, not raw model cleverness.

Its headline number is sharp. On the SemiAnalysis AgentX workload, NVIDIA says Vera Rubin NVL72 systems deliver up to 30x higher throughput per megawatt than GB300 NVL72 on agentic workloads. The same setup also comes with up to 35x lower cost per million tokens than GB300 NVL72. NVIDIA says that translates into 30x more agentic work for the same energy footprint in power-constrained AI factories.

The context matters here. NVIDIA says these workloads are built from real-world agentic coding trajectories, with actual context growth, tool calls and sub-agent spawning preserved. That is a different beast from short chat or summarization jobs, where the token counts usually stay far smaller. In agentic sessions, the model keeps carrying earlier context forward, and every new step feeds the next one.

The company is also selling the supporting cast around the chips. It says DSX MaxLPS can provision up to 40% more GPUs within the same megawatt budget, while NVLink and NVLink Switches in their sixth generation deliver 10x higher packet rates and 3x lower latency than off-the-shelf Ethernet alternatives. NVIDIA also ties the gains to disaggregated serving, rate matching, distributed KV-caching, KV-aware routing, and other tricks that keep long-running sessions from wasting work.

There’s a catch, though. NVIDIA says the Vera Rubin results are early and still pending SemiAnalysis review, and they do not yet include Vera CPU performance for tool calling. Even so, the message is clear: for agents, the scoreboard is moving from tokens per second to tokens per watt. That is a much less glamorous metric, which is exactly why it probably matters more.

My take — AI-written commentary, not fact-checked reporting

This is the kind of announcement that reminds everyone the AI race is really a utility bill race wearing a demo mask. NVIDIA is right to push power efficiency, because nobody wants a datacenter that can reason like a genius and pay like a fire. The industry keeps talking about agents as if intelligence is the hard part; the boring part — energy, latency, and cached context — is where the bill comes due.

Read more about this at: NVIDIA

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.