TLDRocket
Sign in

What it takes to bring up a multi-rack NVIDIA Vera Rubin NVL72 cluster

coreweave.com ● Covered by 6 sources

CoreWeave is bringing up multi-rack NVIDIA Vera Rubin NVL72 clusters. The hard part is making hundreds of GPUs, cooling, power, and networking act like one machine.

Based on reporting by coreweave.com — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

A single NVIDIA Vera Rubin NVL72 rack is already a dense piece of engineering: 72 Rubin GPUs, 36 Vera CPUs, NVLink 6 for scale-up, ConnectX-9 SuperNICs, BlueField-4 DPUs, high-speed storage, power delivery, and 45°C liquid cooling. Once you start linking racks together, the job stops being about one machine and becomes about a distributed AI system with hundreds of GPUs, multiple NVLink domains, thousands of network links, and a lot of ways for one weak spot to slow everything down.

CoreWeave says the only way to keep that under control is to treat the whole stack as software-defined infrastructure from the start. After the hardware is installed and the liquid cooling loop is provisioned, automation takes over: the company detects BMCs, verifies serial numbers, bootstraps credentials, stages nodes, updates firmware, applies metadata, and validates the systems before they move toward production. It has also extended CoreWeave Mission Control with Racky for rack management and Valvey for liquid-cooling control, alongside its Rack LifeCycle Controller, so compute, power, cooling, and environmental sensing all follow the same lifecycle.

The validation process is built to catch the kind of problems that basic diagnostics miss. CoreWeave runs hours of GPU tests, checks CPU-to-GPU data movement, exercises GPU-to-GPU links, pushes compute-heavy workloads to see how the chips handle heat, and then runs realistic training jobs to watch whether each GPU is actually learning correctly. At the rack level, it runs standard and custom benchmarks, tests NVLink behavior, and launches tightly synchronized jobs across all 72 GPUs. Anything that lands below the expected range goes to troubleshooting, not production.

Then the testing widens to the fabric and the cluster. CoreWeave deliberately forces traffic over the backend network, rather than the fast internal GPU links, so it can see how data moves between racks under load. That matters because a link can be up and still be too slow, too uneven, or too flaky for AI work. The company is watching for rising error rates, overheating hardware, uneven traffic, and cabling that does not match the intended map. In a setup like this, a marginal path is enough to leave healthy GPUs waiting around.

Across racks, CoreWeave’s RoCE design uses a two-tier, non-blocking, multi-rail, multi-plane fabric. Each Rubin GPU gets two ConnectX-9 SuperNICs for up to 1.6 Tb/s of backend network connectivity per GPU, and the current modular layout can extend the two-tier architecture to roughly 128,000 GPUs. The whole point is to make the boundary between racks disappear from the workload. If the system is tuned right, customers get one validated pool of Vera Rubin compute instead of a pile of racks that merely happen to be adjacent.

My take — AI-written commentary, not fact-checked reporting

This is the part of AI infrastructure that actually matters: not flashy chip names, but whether the whole pile survives contact with a real workload. The industry loves talking about scale; CoreWeave is talking about the unglamorous stuff that keeps scale from becoming a very expensive outage. That’s the grown-up version of the story, which is rare enough to be refreshing.

Read more about this at: coreweave.com

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.