TLDRocket
Sign in

With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents

NVIDIA NVIDIA Writers Covered by 2 sources

NVIDIA announced that its Vera Rubin NVL72-based rack-scale system, NVIDIA Groq 3 LPX, is now in full production and is being adopted by multiple AI infrastructure partners for agentic inference. In an Artificial Analysis benchmark with Gemma 4 31B on 100,000-token contexts, it delivered 3,400 output tokens per second, about 4x faster than the nearest alternative platform. The update changes the inference stack by pairing long-context context processing on Rubin GPUs with LPX’s lower-latency token decode acceleration, and by extending NVIDIA’s Vera Rubin network fabric to enable larger, lower-cost agent workloads at scale.

Why it matters

The next era of AI inference won’t be defined by a single breakthrough chip, network or system. It’ll be defined by how every layer of the AI factory works together. That’s why NVIDIA is extending Vera Rubin NVL72 with fast token generation for agentic systems. Announced today, the NVIDIA Vera Rubin rack-scale system NVIDIA Groq […]

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.