· cpu / memory
HBM: Trading Data Rate for Bus Width
HBM: Trading Data Rate for Bus Width
DDR5 gains bandwidth by pushing each DQ pin faster. HBM does the opposite: make the bus much wider, make the link much shorter, and keep the per-pin rate modest.
\[\text{bandwidth} = \frac{\text{bus width} \times \text{per-pin rate}}{8}\]A DDR5 channel is 64 bit at 6400 MT/s, about 51 GB/s. An HBM3 stack is 1024 bit at only 6.4 Gb/s, yet delivers 819 GB/s. Same per-pin rate, 16x the width.
1. Structure: stack + TSV + interposer
HBM is not a planar DRAM part but a cube (stack): several DRAM dies stacked vertically, connected by TSVs (through-silicon vias) and microbumps, usually on top of a base die that carries buffer and test logic.
┌──────────────┐ DRAM die (4/8/12/16-Hi)
├──────────────┤ ↕ TSV + microbump
├──────────────┤
├──────────────┤
├──────────────┤ base die (buffer / test logic)
┌────┴──────────────┴────┬──────────────┐
│ silicon interposer │ ← the 1024/2048 wires run here
├────────────────────────┴──────────────┤
│ package substrate │
└───────────────────────────────────────┘
HBM stack GPU / CPU die
The key point: HBM and the processor sit in the same package, connected through a silicon interposer instead of PCB traces. Silicon allows micron-scale pitch, which is what makes a thousand-plus wire bus physically possible. The cost is that the interposer is built with wafer processes, far more expensive than a PCB.
2. Channels and pseudo channels
The HBM interface is split into fully independent channels that are not required to be synchronous with each other.
| Generation | Channel structure | Total width |
|---|---|---|
| HBM1 / HBM2 / HBM2E | 8 channels x 128 bit | 1024 bit |
| HBM3 / HBM3E | 16 channels x 64 bit (2 pseudo channels each, 32 virtual channels) | 1024 bit |
| HBM4 | width doubled | 2048 bit |
Pseudo channels share the command/address bus but have their own data path and bank groups, raising concurrency and bus utilization without adding pins.
3. Generations
| Generation | Standard / year | Width per stack | Per-pin rate | Bandwidth per stack |
|---|---|---|---|---|
| HBM | JESD235, 2013 | 1024 bit | 1 Gb/s | 128 GB/s |
| HBM2 | JESD235A, 2016 | 1024 bit | 2 Gb/s | 256 GB/s |
| HBM2E | HBM2 update, 2018 | 1024 bit | 3.2–3.6 Gb/s | 410–460 GB/s |
| HBM3 | JESD238, 2022 | 1024 bit | 6.4 Gb/s | 819 GB/s |
| HBM3E | vendor implementations, 2023– | 1024 bit | 8–9.6 Gb/s | 1.0–1.2 TB/s |
| HBM4 | JESD270-4, 2025 | 2048 bit | up to 8 Gb/s | up to 2 TB/s |
On capacity: HBM3 supports 4/8/12-Hi stacks (with 16-Hi reserved) and 8Gb–32Gb per layer, giving 4GB to 64GB per stack. HBM3E ships as 8-Hi 24GB and 12-Hi 36GB. HBM4 allows 4–16 layers with 24Gb/32Gb dies, up to 64GB, and stays backward compatible with HBM3 controllers.
4. The signal integrity angle
In the DDR5 posts, reflection, ODT, $R_\mathrm{ON}$ and training take up most of the discussion, because a DDR5 link is long, branched and pluggable: controller → PCB traces → socket → DIMM → multiple DRAM dies.
HBM collapses that into a few millimeters of point-to-point on-silicon interconnect:
DDR5: controller ── PCB (tens to >100 mm) ── socket ── DIMM ── multiple ranks/dies
HBM: controller ── interposer (a few mm) ── one stack
Short link, no stubs, no socket, point-to-point. That in turn enables two things:
- Lower swing: HBM3 uses 0.4V low-swing signaling on the host interface with a 1.1V operating voltage.
- More wires: with less drive and ESD burden per line, running a thousand of them in parallel becomes practical.
This is also why HBM wins on energy per bit — the distance each bit travels is one to two orders of magnitude shorter.
5. What it costs
| Item | Notes |
|---|---|
| Capacity ceiling | Limited stacks per package, far less than a server filled with DIMMs |
| No expansion | Fixed at manufacturing; cannot be upgraded or replaced like a DIMM |
| Cost | Interposer, TSVs, stack bonding, and yield (one bad layer scraps the whole cube) |
| Thermals | Inner dies are trapped in the stack, refresh rate rises with temperature, and it all sits next to a hot GPU |
| Supply | HBM capacity crowds out commodity DRAM wafers, pushing DDR5/NAND prices up |
For reliability, HBM3 adds symbol-based on-die ECC and real-time error reporting — a partial compensation for the fact that the part cannot be replaced.
6. Who uses it
AMD Fiji (2015) was the first HBM GPU. NVIDIA followed with P100 (HBM2), H100 (HBM3), and H200 / Blackwell (HBM3E); AMD Instinct, Intel Xeon Max’s on-package memory, and many AI ASICs and FPGAs use it as well. Volume production today comes from SK hynix, Samsung and Micron.
One notable new direction is SPHBM4 (JESD330-4, 2026): it keeps the HBM4 DRAM stack but narrows the external interface from 2048 bit to 512 bit and compensates with a higher transfer rate, so that the silicon interposer can be dropped in favor of a standard organic substrate.
References
- HBM - Youtube
- High Bandwidth Memory - Wikipedia: per-generation numbers, history, and product adoption.
- JEDEC Publishes HBM3 Update to High Bandwidth Memory (HBM) Standard: 16 channels / 32 pseudo channels, 6.4 Gb/s, 0.4V swing, 1.1V supply, on-die ECC.
- JESD238B.01 High Bandwidth Memory (HBM3) DRAM: the HBM3 standard itself (free download after registration).
-
[High Bandwidth Memory (HBM4) DRAM JEDEC](https://www.jedec.org/standards-documents/docs/jesd270-4): HBM4 standard page. - Micron HBM3E: 8-Hi 24GB, 12-Hi 36GB, >1.2 TB/s product specs.
- Highlights of the High Bandwidth Memory (HBM) Standard: NVIDIA’s Mike O’Connor on HBM vs GDDR bus structure.