TLDRocket
Sign in

CoreWeave brings up multi-rack NVIDIA Vera Rubin NVL72 cluster

coreweave.com ● Covered by 6 sources

CoreWeave turned on a multi-rack NVIDIA Vera Rubin cluster for AI. It also added storage tricks to keep training data closer to the GPUs.

Based on reporting by coreweave.com — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

CoreWeave says it has brought up a multi-rack NVIDIA Vera Rubin NVL72 cluster on CoreWeave Cloud, packing hundreds of Rubin GPUs into one scale-out system for agentic AI. Alongside that, it introduced two new features in CoreWeave AI Object Storage: cross-region write acceleration and a new Archive tier.

The pitch is simple, even if the plumbing is not. Agentic workloads tend to make repeated model calls and tool uses, so delays stack up fast. CoreWeave wants to keep those loops fed by letting jobs write data locally while it copies that data to another region in the background. That way, intermediate state, retrieved context, and outputs do not sit around waiting for cross-region moves.

CoreWeave says it was the first AI cloud provider to validate and bring up a Vera Rubin NVL72. A single rack combines 72 Rubin GPUs with 36 Vera CPUs, NVIDIA NVLink 6, ConnectX-9 SuperNICs, and BlueField-4 DPUs. Multi-rack setups then tie together hundreds of accelerators with NVIDIA Spectrum-X Ethernet into one cluster that can train larger models, serve heavier inference loads, and run reinforcement learning at scale.

The company is also leaning hard on the operational unglamorous bits. Mission Control automates rack setup, hardware detection, firmware updates, validation, power, and cooling. CoreWeave says it validates racks with NVIDIA diagnostics plus full-rack workload testing before production, and it scales the network with modular topology so racks can be added without redesigning everything from scratch.

On the storage side, CoreWeave says its LOTA caching can bring reads down to local NVMe speeds, cut latency by 8x versus a traditional storage cluster, and deliver up to 7 GB/s per GPU. Cécile Robert-Michon of Cohere said the setup helps because their datasets span multiple regions and they cannot let training schedules be ruled by retrieval delays. The new Archive tier is aimed at the stuff teams hate deleting: checkpoints, reproducibility data, and old model versions.

My take — AI-written commentary, not fact-checked reporting

This is the real AI-cloud arms race: not bigger slogans, but fewer excuses for waiting. CoreWeave is betting that whoever hides the storage pain and rack-wrangling best gets the work, which is a more honest strategy than yelling about intelligence. The industry keeps rediscovering that GPUs are only half the machine, and the other half is the boring part people used to ignore.

Read more about this at: coreweave.com

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.