Another Giant Leap: The Rubin CPX Specialized Accelerator & Rack
SemiAnalysis Dylan Patel
Nvidia announced the Rubin CPX, a specialized GPU designed for the prefill phase of AI inference with 20 PFLOPS of FP4 compute but only 2TB/s memory bandwidth, compared to the R200's 33.3 PFLOPS and 20.5TB/s. The Rubin CPX uses 128GB of cheaper GDDR7 memory instead of HBM, reducing memory costs by more than 50% and total production costs significantly. Three new Vera Rubin rack configurations now available in 2026 will combine R200 and Rubin CPX GPUs for disaggregated inference serving, forcing competitors like AMD to redesign their entire roadmaps to develop competing prefill-specialized chips.
Why it matters
Nvidia announced the Rubin CPX, a solution that is specifically designed to be optimized for the prefill phase, with the single-die Rubin CPX heavily emphasizing compute FLOPS over memory bandwidth. This is a game changer for inference, and its significance is surpassed only by the March 2024 announcement of the GB200 NVL72 Oberon rack-scale form […]