With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents
NVIDIA NVIDIA Writers ● Covered by 2 sources
NVIDIA announced that its Vera Rubin NVL72-based rack-scale system, NVIDIA Groq 3 LPX, is now in full production and is being adopted by multiple AI infrastructure partners for agentic inference. In an Artificial Analysis benchmark with Gemma 4 31B on 100,000-token contexts, it delivered 3,400 output tokens per second, about 4x faster than the nearest alternative platform. The update changes the inference stack by pairing long-context context processing on Rubin GPUs with LPX’s lower-latency token decode acceleration, and by extending NVIDIA’s Vera Rubin network fabric to enable larger, lower-cost agent workloads at scale.
Why it matters
The next era of AI inference won’t be defined by a single breakthrough chip, network or system. It’ll be defined by how every layer of the AI factory works together. That’s why NVIDIA is extending Vera Rubin NVL72 with fast token generation for agentic systems. Announced today, the NVIDIA Vera Rubin rack-scale system NVIDIA Groq […]