TLDRocket
Sign in

AI Infra Summit: NVIDIA Vera Rubin and DSX Platform Advancements Showcase Energy Efficiencies of Optimizing Tokens Per Watt for AI Factories

NVIDIA Blog NVIDIA Writers Covered by 2 sources

NVIDIA used the AI Infra Summit to push a simple message: more tokens per watt. That matters because power, not raw speed, is becoming the real AI bottleneck.

Based on reporting by NVIDIA Blog, NVIDIA Writers — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

NVIDIA came to the AI Infra Summit in Santa Clara with a very specific pitch: the next race in AI is no longer just about faster chips, but about squeezing more useful work out of every megawatt. Ian Buck, NVIDIA’s vice president of hyperscale and high-performance computing, used a packed session to show how the company is tying hardware, networking, software and power management together into one factory-style system.

The crowd was noticeably bigger too. NVIDIA said the summit drew more than 8,000 attendees this year, up from 3,500 last year. That kind of growth fits the mood of the event, which has clearly moved from niche infrastructure meetup to something much louder. And the announcements all pointed in the same direction: power is the constraint now.

NVIDIA’s partners were front and center. Amazon’s Annapurna Labs is working with NVIDIA on NVHBM custom high-bandwidth memory. d-Matrix is integrating with NVLink Fusion to combine NVIDIA Vera CPUs with d-Matrix Raptor XPUs for low-latency inference at scale. Emerald AI and NVIDIA also showed a commercial flexible-load program with Silicon Valley Power, while Lambda reported a 23% improvement in performance per watt using NVIDIA DSX MaxLPS. Pinterest, meanwhile, is using Blackwell and NVIDIA Dynamo inference software for conversational AI in visual discovery.

The underlying strategy is bigger than any one deal. NVIDIA is pushing a full-stack AI factory platform built around Vera Rubin systems, Dynamo, NeMo, and its networking gear, including NVLink, Spectrum-X Ethernet, ConnectX SuperNICs, BlueField-powered storage and BlueField DPUs. The pitch is that AI factories should be codesigned from silicon to grid, and that the real metric is shifting from peak performance to validated agentic tokens per megawatt.

The numbers NVIDIA chose were aggressive. DSX MaxLPS is claimed to deliver up to 1.4x more tokens per megawatt through factory-wide power optimization. Lambda said it could run 19 nodes in the same power budget usually reserved for 16 full-power nodes, raising cluster-wide token throughput by 24% to about 5 million tokens per second while lifting performance per watt by 23%. NVIDIA also said Vera Rubin NVL72, combined with Groq 3 LPX, can reach up to 35x higher token throughput per megawatt than GB200 NVL72 for large models at long context.

What makes the pitch land is that NVIDIA isn’t talking about this as theory. Emerald AI’s system responded to hundreds of demand signals from Silicon Valley Power while protecting workload performance. Vera Rubin NVL72 results on SemiAnalysis AgentX showed up to 30x higher throughput per megawatt than GB300 NVL72 on DeepSeek V4 Pro. Reliability got its own spotlight too, with NVLink 6 framed as a way to keep massive AI factories running without faults rippling outward. The message is blunt: if power is scarce, the winners will be the companies that turn watts into tokens with the least waste.

My take — AI-written commentary, not fact-checked reporting

NVIDIA has correctly spotted the new obsession: not raw speed, but how much work a data center can squeeze out of the wall socket. That’s a grown-up pitch, and it beats the usual chip-bro fireworks by a mile. The catch is obvious: once everyone starts selling “efficiency” as the headline, the real competition moves to who can prove it under pressure, not just on stage.

Read more about this at: NVIDIA Blog

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.