Why AI accelerators are wrapped in stacks of HBM
Open any photo of a modern AI GPU and you'll see the giant compute die in the middle, ringed by short, fat towers of memory soldered millimeters away. Those towers are HBM, and they exist because regular DRAM physically cannot feed a matrix engine fast enough.
On this page
The picture version
Six pictures for a reader who has never thought about where a chip keeps its numbers. The prose below fills in the seams the pictures skip.
1 · The problem
The chef works in seconds. The pantry is two blocks away.
2 · The obvious fix, and the wall it hits
Send the bits faster. Physics on a circuit board disagrees.
3 · The other dial
If you can’t run faster, run wider. Very much wider.
4 · Where a thousand wires can actually fit
Build the floor out of silicon, and the memory has to move next door.
5 · Out of floor space
If you can’t grow outward, grow up.
6 · Keep this card
The whole thing on one index card.
Why it exists
You’ve watched an AI chatbot type its answer out at you, word by word, at a pace you can comfortably read along with. That’s odd when you think about it: the chip generating those words can do trillions of arithmetic operations per second, and a sentence is a few dozen words. It isn’t thinking that slowly. It’s waiting — for the model’s weights to arrive from memory, over and over, once per word.
Picture a busy restaurant where the chef can cook a dish in 5 seconds, but the pantry is in a building two blocks away. It doesn’t matter how fast the chef is — the kitchen runs at the speed of the runner fetching ingredients. To go faster, you don’t hire a faster chef. You move the pantry into the kitchen. High Bandwidth Memory — HBM — is the AI-chip version of that move. Instead of making the GPU compute faster, designers physically picked up the memory chips and stacked them millimeters away from the processor, on the same package. The chef and the pantry now share a counter. That short distance is a large part of why frontier models can be served at all.
Look at a die-shot of a GPU meant for AI — an H100, an MI300, a TPU package — and the visual is almost always the same. There’s a big square of compute logic in the middle, and right up against it, separated by a millimeter or two of silicon interposer, sit four to eight short rectangular towers. Those towers are the HBM stacks. They look out of place — like someone glued extra chips to the GPU — and that visual oddity is the whole story.
Regular computer memory doesn’t sit there. DDR sticks live a few centimeters away on the motherboard, connected by long copper traces. GDDR chips, used in gaming GPUs, sit a centimeter away on the same PCB. HBM sits next to the compute die on the same package, glued in by a special interposer, with thousands of wires running between them.
The reason is brutally simple: AI workloads don’t need more compute as much as they need more bandwidth — bytes per second from memory into the matrix multiplier — and the only way to get that many bytes that fast is to put the memory practically inside the chip.
Why it matters now
That chatbot typing at you is the everyday version of a much larger pattern: a lot of frontier-model serving is a memory-bandwidth problem dressed up as a compute problem. Generating one token for one user means reading the model’s weights and the KV cache out of memory and running them through the matrix multipliers exactly once. In that regime the multipliers are barely the bottleneck; the wires between memory and multipliers are. (Batch enough users together, or switch to training, and the balance shifts back toward compute — the bottleneck is a property of the workload, not of the chip.) This is why a chip’s HBM bandwidth — measured in terabytes per second, not gigabytes — ends up being one of the most-quoted numbers in any new AI silicon launch, sometimes ahead of the FLOP count.
It’s also why chip supply has gotten weird. The HBM market is dominated by a handful of DRAM manufacturers (publicly: SK hynix, Samsung, Micron), and constraints there now constrain who can build AI accelerators at all. The compute logic isn’t usually the gating part. The stacked memory is.
The short answer
HBM = stacked DRAM dies + through-silicon vias + silicon interposer + very wide, slow bus
Picture to keep: not a faster pipe — a wider one. HBM’s per-pin rate is unremarkable — comparable to a DDR5 stick, slower than GDDR; what it does is run a thousand wires side by side over a distance short enough that a thousand wires is physically possible.
HBM is just regular DRAM cells, rearranged. You take eight or twelve DRAM dies, stack them physically on top of each other, drill vertical wires straight through the silicon to connect them (those are TSVs), and then sit that whole tower on an interposer right next to the compute die. The bus to the compute die is deliberately slow per wire — HBM’s per-pin data rates are lower than GDDR’s, not higher — but the bus is very wide, often 1024 bits per stack. Bandwidth = width × rate per wire, and when you can’t push the rate further, you push the width.
How it works
Follow the chef-and-pantry problem through, one failed fix at a time.
Naive attempt: use normal memory. Bolt DDR sticks to the motherboard next to the accelerator, the way every server has done for decades. This fails immediately and quantitatively — the matrix units can consume bytes far faster than a handful of DDR channels can deliver them, so the chip spends most of its time idle. The runner is too slow, and the kitchen runs at the runner’s speed.
Fix 1: make the memory faster. Push the clock up.
Memory bandwidth for any DRAM technology is roughly bus width in bits × data rate per pin. Push the per-pin rate too hard and signal integrity falls apart on the long PCB traces — DDR5 sticks max out around 6–8 GT/s per pin in practice, GDDR pushes higher because the traces are shorter, but every step up the clock costs more power for diminishing returns. The physics is set by capacitance, inductance, and how loudly a wire couples to its neighbors. You can’t just print “10 GHz” on the box.
Fix 2: run wider instead. If you can’t run faster, run more wires in parallel. A DDR5 channel is 64 bits. A modern HBM stack is 1024 bits (split into 16 channels of 64 bits each). Instead of trying to win the GHz race, HBM wins the parallel-pins race. An H100-class part, depending on the SKU, carries five or six HBM stacks and moves a few terabytes per second. The same compute die paired with sticks of DDR would move a fraction of that.
But you can’t have 1024 wires per stack on a normal PCB.
This is where the interposer matters. A printed circuit board can route maybe a few hundred fast signals to a chip before you run out of layers and space. A silicon interposer — basically a thin extra slab of silicon underneath both the compute die and the HBM stacks — is fabricated with the same lithography used for chips, so it can carry thousands of microscopic traces packed tightly together. That’s the only way a 1024-bit-per-stack bus is physically buildable.
The interposer is also why HBM has to live so close. Those traces are microscopic, and their reach is short: push the distance and you start paying in signal integrity and power. Exactly how short depends on the signalling and packaging, and no clean threshold is published — but it’s millimetres, not centimetres, which is why HBM towers ring the compute die. They have nowhere else to go.
And now the memory has nowhere to live. Ringing the compute die with memory that must stay within a few millimetres leaves you very little floor space — but AI workloads want tens of gigabytes. That’s what forces the last move: if you can’t grow outward, grow up.
A single DRAM die only stores so many bits. To hit the tens of gigabytes per stack that AI workloads want, the dies are physically stacked — eight high, twelve high, sometimes more — and connected vertically through TSVs. Stacking trades cost and yield (you have to bond and test each layer; one bad die can ruin the stack) for density and very short vertical wires.
The standard account is that this stacking is genuinely hard manufacturing. Yield on a 12-high stack is worse than yield on a single die, and thermals are awkward — heat from the bottom die has to leave through layers of memory above it. The major manufacturers don’t publish yield numbers, so take “it’s hard” as the qualitative claim, not a specific figure.
Why this favors AI workloads specifically.
Go back to the chatbot. Generating one token for one user reads gigabytes of weights, runs each one through a multiply exactly once, and moves on. The arithmetic-to-memory ratio (the “arithmetic intensity”) is about as low as it gets — each weight is read and used once. That puts the workload on the memory-bandwidth side of the roofline, and HBM exists for exactly that regime.
The honest qualifier: not every transformer kernel lives there. Batch many users together, or train instead of serve, and the same weights get reused across many inputs, intensity climbs, and the kernel moves toward the compute ceiling. Training also drags in its own memory pressure (gradients, activations, optimizer state). So “bandwidth-bound” is a claim about a workload, not a law about transformers — but the workload that dominates interactive serving is squarely in HBM’s regime. And it’s still overkill for a laptop CPU running a database, where caches and a couple of DDR channels keep up fine.
The pantry-in-the-kitchen picture is right about distance and wrong about capacity: a real pantry gets bigger when you need more food, and HBM can’t. It’s stuck with whatever fits in a few stacks beside the die, which is why a model that doesn’t fit in HBM is a much bigger problem than a model that runs slowly — you fall off a cliff to system memory rather than sliding down a slope.
And that answers the die-shot: the towers are glued to the edge of the compute die because they have to be. A thousand-wire bus can only be built out of chip-scale traces, and chip-scale traces only reach a few millimetres. The layout isn’t a packaging choice — it’s the bus width made visible.
You started with HBM = stacked DRAM dies + TSVs + interposer + very wide, slow bus. What did this post add? — + "slow" is not a compromise, it's the design. Every part of HBM exists to make a wide bus physically possible, because the clock-speed lever was already pulled to the end. Stacking, TSVs, and the interposer aren’t three features; they’re three consequences of one decision to buy bandwidth by the wire instead of by the hertz.
Check yourself
Before you go — a vendor announces a new accelerator with 2× the FLOPs of the last one and the same HBM bandwidth. For a single-user LLM chat session, roughly how much faster should you expect token generation to be?
Answer
Barely faster, if at all. Generating one token for one user means streaming the entire weight set out of memory and multiplying it against a single vector — arithmetic intensity is low, so the kernel is sitting on the memory-bandwidth slope of the roofline, not under the compute ceiling. Doubling the ceiling doesn’t move a workload that never touches it. You’d expect a real speedup on training or on large-batch serving, where intensity is high enough to be compute-bound, and little on single-stream decoding. (Some gain can still leak in from better caches, fused kernels, or lower-precision formats that also cut bytes moved — but the FLOP number alone doesn’t buy it.) This is exactly why bandwidth numbers get quoted alongside, and sometimes ahead of, FLOP numbers in launch decks.
And one more — if wider buses are so much better than faster clocks, why doesn’t your laptop CPU use HBM?
Answer
Because the interposer is the expensive part and a CPU doesn’t need it. HBM’s width only pays off if your workload is memory-bandwidth bound; a CPU running a browser, a compiler, or a database has high enough cache hit rates and low enough streaming demand that a couple of DDR channels keep it fed. What you’d be buying instead is packaging cost, worse yield (a bad DRAM die can spoil a whole stack), harder thermals, and — critically — a memory capacity fixed at manufacture, since you can’t slot in more HBM the way you can add a DIMM. For general-purpose computing that’s a bad trade in every dimension.
Famous related terms
- GDDR —
GDDR ≈ DDR + shorter traces + more aggressive clocks— what gaming GPUs use; bandwidth between DDR and HBM, but cheap and on a normal PCB. - DDR / LPDDR —
DDR = standard PC main memory + 64-bit channels + commodity sticks— what your laptop and server CPU talk to. Plenty for general computing, far too narrow for a matrix engine. - Through-silicon via (TSV) —
TSV = vertical hole + plated metal + die-to-die connection— the physical trick that makes a stack act like one chip electrically. - Silicon interposer —
interposer = thin silicon slab + microscopic traces + a way to wire two dies together at chip density— what lets HBM sit next to the compute die instead of across a motherboard. - Memory bandwidth — see memory-bandwidth — the metric HBM exists to maximize.
- Roofline model —
roofline = peak FLOPs ceiling + memory-bandwidth slope— the chart that says, for a given kernel, whether compute or memory is your wall.
Going deeper
- The JEDEC HBM standards (HBM through HBM3E and later) — for “what are the actual bus widths, stack heights, and per-pin rates?”, the only place those numbers are defined rather than marketed. Free but registration-walled, which is why most people quote them second-hand.
- AMD and SK hynix’s 2015 “High Bandwidth Memory: The New Standard for Graphics” — for “why does the interposer-and-stack design follow from wanting bandwidth?”, explained by the people who shipped the first HBM part, before AI was the justification.
- Rabbit hole: Williams, Waterman, and Patterson, “Roofline: An Insightful Visual Performance Model” — for “would my kernel even benefit from more bandwidth?”, which is the question that decides whether any of this matters to you.
One gap worth naming: there is no reliable public source for current HBM stack yields or per-stack manufacturing costs. The figures in circulation come from analyst firms (TrendForce, SemiAnalysis) rather than from the manufacturers, who don’t publish them.