Scaling the Memory Wall: The Rise and Roadmap of HBM
SemiAnalysis Dylan Patel
High Bandwidth Memory (HBM) has become essential for AI accelerators, with demand driven by AI model scaling and accelerator designs from Nvidia, Google, OpenAI, and others. The report projects Nvidia will command the largest share of HBM bit demand in 2027, with its Rubin Ultra pushing per-GPU capacity to 1 TB, while demand grows across custom ASICs and hyperscalers like Amazon and Google. Manufacturing challenges including power delivery network design, thermal management, and bonding tool precision create yield pressures that differentiate suppliers, with SK Hynix and Micron achieving higher yields than Samsung, thereby sustaining premium HBM pricing and margin growth.
Why it matters
The first portion of this report will explain HBM, the manufacturing process, dynamics between vendors, KVCache offload, disaggregated prefill decode, and wide / high-rank EP. The rest of the report will dive deeply into the future of HBM. We will cover the revolutionary change coming to HBM4 with custom base dies for HBM, what various […]