How XPUs Meet a World-Class AI Factory
NVIDIA Jesse Clayton
NVIDIA is pitching NVLink Fusion to plug custom XPUs into its AI factory stack. The pitch: faster launches, less risk, and better performance than wiring everything yourself.
Based on reporting by NVIDIA, Jesse Clayton — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
NVIDIA is making a simple argument with a lot of engineering behind it: if AI is going to be built like a factory, then custom XPUs should not have to reinvent the whole factory to get to market. The company’s new NVLink Fusion program is meant to connect those chips to NVIDIA’s own AI infrastructure, so builders can keep the custom parts custom and borrow the rest from a system that already exists.
That matters because the source of truth in an AI factory is not raw chip count. It is tokens per second, tokens per watt, cost per token, utilization, and uptime. Once the scale gets large enough, a weak scale-up fabric can push utilization down and costs up. NVIDIA says NVLink Fusion is aimed at workloads like trillion-parameter models, mixture-of-experts systems, and agentic AI, where performance, resiliency, and platform maturity all have to hold together at once.
The headline hardware claim is a 72-XPU NVLink domain built around sixth-generation NVLink. NVIDIA says XPU-to-XPU latency is 3x lower than alternatives based on off-the-shelf Ethernet, while packet rate is 10x higher. The company also points to GB300 NVL72 systems as delivering higher throughput and better interactivity than setups that do not use NVL72, and says future NVLink roadmaps will stretch to domains of up to 1,152 accelerators plus co-packaged optics.
NVLink Fusion also reaches beyond the accelerator itself. It includes NVLink-C2C for linking XPUs to NVIDIA Vera CPUs or other ecosystem CPUs, with up to 6x the energy efficiency of PCIe. That is meant to narrow the gap between control and compute in agentic systems, while keeping the rest of the platform on familiar ground.
And that familiar ground is doing a lot of work here. NVIDIA and its partners frame the real challenge as everything around the chip: CPU and scale-up interfaces, network validation, rack design, cooling, power, security, storage, supplier coordination, and serviceability. The program ties into NVIDIA MGX rack-scale architecture, the NVIDIA DSX reference architecture, and the Omniverse DSX AI Factory Blueprint, which is pitched as a digital twin and open reference design for gigawatt-scale AI factories.
The software stack is part of the pitch too, with NCCL for distributed workloads, Dynamo and NIXL for disaggregation, and Mission Control for cluster management, telemetry, and debugging. The whole idea is to let hyperscalers and AI-native companies move ahead with buildout before the final silicon mix is settled, then shift capacity later as workload demand, supply, and business priorities change.
My take — AI-written commentary, not fact-checked reporting
This is NVIDIA doing what NVIDIA does best: turning dependency into a platform and calling it flexibility. Custom chip makers love the word semi-custom right up until they have to build the cooling, racks, software, and supply chain that make the chip useful. The more AI starts to look like infrastructure, the more the boring parts win.
Read more about this at: NVIDIA
Related stories
NVIDIA NVLink Fusion Expands With NVHBM Custom High-Bandwidth Memory
NVIDIA · 1 week ago ·
21
Nvidia and Cisco push the enterprise AI factory into the rack-scale era
SiliconANGLE · 1 week ago ·
45
AI Innovators Adopt NVIDIA Vera — Why Max Single-Threaded CPU at Scale Matters
NVIDIA · 1 month ago ·
30