Learning Record learning from practice

· cpu / memory

HBM: Trading Data Rate for Bus Width

HBM: Trading Data Rate for Bus Width

DDR5 gains bandwidth by pushing each DQ pin faster. HBM does the opposite: make the bus much wider, make the link much shorter, and keep the per-pin rate modest.

\[\text{bandwidth} = \frac{\text{bus width} \times \text{per-pin rate}}{8}\]

A DDR5 channel is 64 bit at 6400 MT/s, about 51 GB/s. An HBM3 stack is 1024 bit at only 6.4 Gb/s, yet delivers 819 GB/s. Same per-pin rate, 16x the width.

1. Structure: stack + TSV + interposer

HBM is not a planar DRAM part but a cube (stack): several DRAM dies stacked vertically, connected by TSVs (through-silicon vias) and microbumps, usually on top of a base die that carries buffer and test logic.

        ┌──────────────┐  DRAM die (4/8/12/16-Hi)
        ├──────────────┤   ↕ TSV + microbump
        ├──────────────┤
        ├──────────────┤
        ├──────────────┤  base die (buffer / test logic)
   ┌────┴──────────────┴────┬──────────────┐
   │        silicon interposer             │  ← the 1024/2048 wires run here
   ├────────────────────────┴──────────────┤
   │            package substrate          │
   └───────────────────────────────────────┘
            HBM stack        GPU / CPU die

The key point: HBM and the processor sit in the same package, connected through a silicon interposer instead of PCB traces. Silicon allows micron-scale pitch, which is what makes a thousand-plus wire bus physically possible. The cost is that the interposer is built with wafer processes, far more expensive than a PCB.

2. Channels and pseudo channels

The HBM interface is split into fully independent channels that are not required to be synchronous with each other.

Generation Channel structure Total width
HBM1 / HBM2 / HBM2E 8 channels x 128 bit 1024 bit
HBM3 / HBM3E 16 channels x 64 bit (2 pseudo channels each, 32 virtual channels) 1024 bit
HBM4 width doubled 2048 bit

Pseudo channels share the command/address bus but have their own data path and bank groups, raising concurrency and bus utilization without adding pins.

3. Generations

Generation Standard / year Width per stack Per-pin rate Bandwidth per stack
HBM JESD235, 2013 1024 bit 1 Gb/s 128 GB/s
HBM2 JESD235A, 2016 1024 bit 2 Gb/s 256 GB/s
HBM2E HBM2 update, 2018 1024 bit 3.2–3.6 Gb/s 410–460 GB/s
HBM3 JESD238, 2022 1024 bit 6.4 Gb/s 819 GB/s
HBM3E vendor implementations, 2023– 1024 bit 8–9.6 Gb/s 1.0–1.2 TB/s
HBM4 JESD270-4, 2025 2048 bit up to 8 Gb/s up to 2 TB/s

On capacity: HBM3 supports 4/8/12-Hi stacks (with 16-Hi reserved) and 8Gb–32Gb per layer, giving 4GB to 64GB per stack. HBM3E ships as 8-Hi 24GB and 12-Hi 36GB. HBM4 allows 4–16 layers with 24Gb/32Gb dies, up to 64GB, and stays backward compatible with HBM3 controllers.

4. The signal integrity angle

In the DDR5 posts, reflection, ODT, $R_\mathrm{ON}$ and training take up most of the discussion, because a DDR5 link is long, branched and pluggable: controller → PCB traces → socket → DIMM → multiple DRAM dies.

HBM collapses that into a few millimeters of point-to-point on-silicon interconnect:

DDR5:  controller ── PCB (tens to >100 mm) ── socket ── DIMM ── multiple ranks/dies
HBM:   controller ── interposer (a few mm) ── one stack

Short link, no stubs, no socket, point-to-point. That in turn enables two things:

  • Lower swing: HBM3 uses 0.4V low-swing signaling on the host interface with a 1.1V operating voltage.
  • More wires: with less drive and ESD burden per line, running a thousand of them in parallel becomes practical.

This is also why HBM wins on energy per bit — the distance each bit travels is one to two orders of magnitude shorter.

5. What it costs

Item Notes
Capacity ceiling Limited stacks per package, far less than a server filled with DIMMs
No expansion Fixed at manufacturing; cannot be upgraded or replaced like a DIMM
Cost Interposer, TSVs, stack bonding, and yield (one bad layer scraps the whole cube)
Thermals Inner dies are trapped in the stack, refresh rate rises with temperature, and it all sits next to a hot GPU
Supply HBM capacity crowds out commodity DRAM wafers, pushing DDR5/NAND prices up

For reliability, HBM3 adds symbol-based on-die ECC and real-time error reporting — a partial compensation for the fact that the part cannot be replaced.

6. Who uses it

AMD Fiji (2015) was the first HBM GPU. NVIDIA followed with P100 (HBM2), H100 (HBM3), and H200 / Blackwell (HBM3E); AMD Instinct, Intel Xeon Max’s on-package memory, and many AI ASICs and FPGAs use it as well. Volume production today comes from SK hynix, Samsung and Micron.

One notable new direction is SPHBM4 (JESD330-4, 2026): it keeps the HBM4 DRAM stack but narrows the external interface from 2048 bit to 512 bit and compensates with a higher transfer rate, so that the silicon interposer can be dropped in favor of a standard organic substrate.

References