TLDRocket
Sign in

From Training to Production, NVIDIA and CoreWeave Close the Loop on Agentic AI

NVIDIA Blog Stuart Pitts

NVIDIA and CoreWeave are putting Vera Rubin into real production, not just demos. Cognition is already using it, and CoreWeave says the new setup speeds up agent work and model tuning.

Based on reporting by NVIDIA Blog, Stuart Pitts — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

CoreWeave and NVIDIA are pushing their long partnership into a new phase: production. This week at CoreWeave Fully Connected in San Francisco, CoreWeave said NVIDIA Vera Rubin NVL72 systems with Spectrum-X 102.4T Ethernet are now available on its cloud. Cognition, the lab behind the Devin AI software engineer, is the first customer running production workloads on Vera Rubin.

That matters because the pitch here is not just faster chips, but a tighter loop between training and deployment. CoreWeave also plans to offer NVIDIA Vera, described as the first CPU built for AI agents, and it launched CoreWeave Forge, a connected environment for training, evaluating and improving models and agents on NVIDIA accelerated computing.

The performance claims are pointed. Cognition benchmarked Vera Rubin NVL72 against a GB200 NVL72 baseline using a real software engineering workload built from a subset of tasks sampled from FrontierCode. In early tests, it saw up to 4.8x higher total token throughput for SWE-2 inference workloads. For Devin, that translates into faster code generation and quicker multistep reasoning.

CoreWeave says it stood up a production Vera Rubin cluster for Cognition in days, helped by co-design across the stack. On the CPU side, testing of NVIDIA Vera showed more than 3x faster agent sandbox startup times, and a 1.7x performance gain on Terminal-Bench across all passing tasks. CoreWeave says a single rack can hold 128 CPUs and 11,264 cores, enough for more than 11,000 concurrent environments at one core each.

Forge is the other half of the story. It combines Weights & Biases, OpenPipe and marimo in one environment, and adds tools like CoreWeave ARIA, now generally available, and Agent Lens, a new service for production observability. CoreWeave says Agent Lens improves failure detection by 20% and cuts fix costs in half. The company also says its serverless RL trains 1.4x faster at 40% lower cost than a self-managed setup, while NVIDIA Dynamo powers its managed inference service and RL Rollouts is in private preview.

My take — AI-written commentary, not fact-checked reporting

This is the unglamorous part of AI that actually matters: plumbing, not posters. The winners will be the companies that can keep models, agents and production traces in one loop without making teams play courier between vendors. The rest is just expensive theater with better GPUs.

Read more about this at: NVIDIA Blog

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.