NVIDIA Vera Rubin Driving Performance Per Watt, Lowest Token Cost for Partners Worldwide
NVIDIA NVIDIA Writers ● Covered by 8 sources
Nvidia's Vera Rubin AI chip platform is now in production at CoreWeave, Google, Microsoft and Oracle. Early benchmarks claim 10x more AI output per watt than the previous Blackwell generation.
Based on reporting by NVIDIA, NVIDIA Writers — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Nvidia doesn't do quiet product launches, and Vera Rubin is no exception. The company says its next-generation AI rack system, the NVL72, is now ramping into production across CoreWeave, Google Cloud, Microsoft Azure and Oracle Cloud Infrastructure, backed by a supply chain spanning more than 350 factories in 30 countries. That's an enormous logistical claim, but the headline number is the one that matters to anyone paying electricity bills: CoreWeave's first DeepSeek-R1 benchmark on Vera Rubin showed 10x more throughput per megawatt than the Grace Blackwell NVL72 it replaces.
The gains come from what Nvidia calls extreme co-design — building seven chips and five rack trays as one integrated system rather than bolting together off-the-shelf parts. The new Vera CPU sits at the center of this, built specifically for agentic workloads with a custom Olympus core that doubles single-threaded performance and cuts memory latency by 40% compared to rival chiplet designs. Networking gets the same treatment: sixth-generation NVLink promises over twice the throughput on complex jobs, and Spectrum-X Ethernet with new Spectrum-6 switches claims 1.6x higher bandwidth than standard Ethernet setups. Nvidia also touts its co-packaged optics switch, already in volume production and adopted early by CoreWeave, Lambda and OCI, as a way to cut power use fivefold versus pluggable transceivers.
The more interesting story, though, might be in Europe. Microsoft and Mistral are expanding their partnership on the back of Vera Rubin, with a new multibillion-dollar infrastructure deal that gives Mistral access to thousands of Rubin GPUs. The pitch is sovereignty without sacrifice — European governments and regulated industries get frontier-grade AI models like Mistral Medium 3.5, deployable across public cloud, private cloud or fully disconnected environments, without having to depend on infrastructure outside their control. Nvidia claims Vera Rubin NVL72 delivers 10x more tokens per megawatt and a tenth of the cost per million tokens compared to the older GB200 NVL72, numbers clearly aimed at governments wary of both American tech dependency and runaway compute bills.
Elsewhere, the rollout reads like a checklist of early adopters proving out different workloads. Google Cloud's new A5X instance, running on Vera Rubin, is already powering London startup Ineffable Intelligence, which is building reinforcement-learning
Read more about this at: NVIDIA