CoreWeave and NVIDIA enable Cognition to run production workloads on NVIDIA Vera Rubin NVL72 and launch CoreWeave Forge to feed observability back into training and evaluation
Deployment ● Confirmed 74% confidence first seen
CoreWeave announced that Cognition can run production workloads on NVIDIA Vera Rubin NVL72, with Vera Rubin availability and NVIDIA Vera CPU support provided through CoreWeave Cloud. The coverage also describes CoreWeave Forge, which connects production observability back into training and evaluation, supporting an iterative “training-to-production” workflow and reporting throughput gains for SWE-2 inference on Vera Rubin versus a GB200 NVL72 baseline.
Decision brief
- What changed
- CoreWeave announced that Cognition can run production workloads on NVIDIA Vera Rubin NVL72, and that Vera Rubin availability plus NVIDIA Vera CPU support will be offered through CoreWeave Cloud. CoreWeave also launched CoreWeave Forge, a platform intended to feed production observability back into training and evaluation for iterative model and agent improvement.
- Why it matters
- This matters because it links next-generation inference infrastructure with a workflow for using live production signals to improve models faster, which can affect how companies organize AI deployment, evaluation, and infrastructure spend. For leaders operating agentic or always-on AI systems, the combination of claimed higher throughput on Vera Rubin and a closed-loop training-to-production toolchain could influence platform selection, capacity planning, and MLOps architecture assumptions. The practical implication, assuming these capabilities generalize beyond the showcased workload, is a tighter connection between production reliability, model iteration speed, and unit economics.
- Evidence
- The core facts come primarily from NVIDIA’s own coverage of the announcement and product positioning, including the claim that Cognition saw up to a 4.8x token-throughput increase for SWE-2 inference on Vera Rubin NVL72 versus a GB200 NVL72 baseline. SiliconANGLE’s separate reporting is directionally consistent on the continuous-learning workflow, Cognition use case, and AI-factory economics focus, but the coverage remains largely based on company statements rather than independent benchmarking.
- What remains uncertain
- It is still unclear how broadly the reported throughput gains apply across other models, workloads, latency targets, and total cost profiles, since the cited performance is tied to Cognition’s SWE-2 inference and NVIDIA/CoreWeave framing. The coverage does not independently verify Forge’s operational impact, pricing, deployment complexity, or whether most enterprises can reproduce the same closed-loop training and production workflow at scale.
- Monitor next
- Watch for independent customer deployments or benchmark disclosures showing Vera Rubin NVL72 production performance and measurable business outcomes from CoreWeave Forge outside the Cognition example.
Analytical support, not advice — assumptions and open questions stated above.