NVIDIA announced improved efficiency and production readiness for Vera Rubin NVL72-based systems used to run agentic long-context inference workloads
Product launch Provisional 70% confidence first seen
NVIDIA reported that Vera Rubin NVL72 systems can deliver significantly higher agentic-workload throughput per watt than prior configurations, lowering estimated cost per million tokens based on the SemiAnalysis AgentX workload. The company also said its Vera Rubin NVL72-based NVIDIA Groq 3 LPX rack-scale platform is now in full production and benchmarked higher output token throughput on long-context agent inference.