Skip to main content
Back to Glossary
ai-infrastructure

High Bandwidth Memory

HBM

A 3D-stacked DRAM standard co-packaged on the GPU die, delivering memory bandwidths of 3–5 TB/s to feed the parallel compute engines of modern AI accelerators.

Unit: TB/s (bandwidth) / GB (capacity)

Full Definition

High Bandwidth Memory (HBM) uses through-silicon vias (TSVs) to stack multiple DRAM dies vertically, then co-packages the stack with the GPU using 2.5D interposer technology (TSMC CoWoS or equivalent). This proximity gives HBM bandwidths of 3.35 TB/s (HBM3) to 4.8 TB/s (HBM3e) — versus 100–200 GB/s for conventional GDDR6 — at lower power per bit transferred.

AI model training is frequently memory-bandwidth-bound rather than compute-bound; HBM directly addresses this bottleneck. NVIDIA H100 SXM5 incorporates 80 GB of HBM3 at 3.35 TB/s; the H200 upgrades to 141 GB of HBM3e at 4.8 TB/s. HBM's high cost and limited supply are significant factors in GPU platform economics and datacenter procurement lead times.

Also Known As

HBM2eHBM3HBM3estacked memory

Source Reference

JEDEC JESD235 HBM Standard; NVIDIA H200 Tensor Core GPU Datasheet

Related Terms