NVIDIA NVLink Fusion Expands With NVHBM Custom High-Bandwidth Memory
NVIDIA Jesse Clayton
NVIDIA added NVHBM to NVLink Fusion for custom AI chips. It promises more bandwidth, less power, and easier memory integration for hyperscalers.
Based on reporting by NVIDIA, Jesse Clayton — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
NVIDIA has widened its NVLink Fusion push with a new memory option called NVHBM, aimed at the kind of AI systems that are starting to strain every part of the stack at once. The pitch is simple: when trillion-parameter workloads and AI agents get bigger, memory can’t be an afterthought anymore.
NVHBM is a next-generation high-bandwidth memory technology meant for XPUs, and NVIDIA says it will be validated and sold through leading memory partners. The key change is architectural. Instead of putting the memory controller on the XPU die, NVHBM moves NVIDIA’s custom controller into the HBM base die.
That shift is supposed to matter in three ways. NVIDIA says NVHBM can deliver up to 30% more memory bandwidth, cut HBM power consumption by 15%, and free up to 25% more area on the XPU compute die compared with standard HBM4E. The company also says it is setting up a standard NVHBM implementation that multiple memory providers can offer, which should reduce the work needed to integrate and qualify memory from different suppliers.
Amazon’s Annapurna Labs is the first partner named for the technology. It will work with NVIDIA on NVHBM and on the NVLink scale-up architecture, building on AWS’s earlier support for NVLink Fusion. Amazon said its next-generation Trainium chips, starting with Trainium4, will support NVLink Fusion so Amazon chips and NVIDIA GPUs can operate together in a common rack-scale architecture.
NVLink Fusion itself is NVIDIA’s way of letting partners connect custom XPUs and CPUs to its rack-scale platform. The company says partners can use NVIDIA NVLink chiplets, NVLink-C2C, NVLink Switches and MGX systems and racks, plus a wider ecosystem of CPU, ASIC, system and technology vendors. The selling point is not subtle: less custom integration pain, less risk, and a quicker path to semi-custom AI infrastructure.
My take — AI-written commentary, not fact-checked reporting
NVIDIA is doing what NVIDIA does best: turning a bottleneck into a product and then calling it an ecosystem. The real story here isn’t just faster memory, it’s the growing pressure on hyperscalers to accept NVIDIA’s rack-scale rules if they want their custom chips to behave nicely. That’s open in the sense that others can join, but only after they agree to play on NVIDIA’s court.
Read more about this at: NVIDIA