Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why AI accelerators are wrapped in stacks of HBM

Open any photo of a modern AI GPU and you'll see the giant compute die in the middle, ringed by short, fat towers of memory soldered millimeters away. Those towers are HBM, and they exist because regular DRAM physically cannot feed a matrix engine fast enough.

Science intermediate Apr 29, 2026 · updated Aug 25, 2026 · 13 min read

On this page

The picture version

Six pictures for a reader who has never thought about where a chip keeps its numbers. The prose below fills in the seams the pictures skip.

1 · The problem

The chef works in seconds. The pantry is two blocks away.

the chip trillions of sums a second and mostly idle “send me the numbers” …a long wait… the memory out on the circuit board A faster chef doesn’t help. The bottleneck is the walk to the pantry. which is why a chatbot types at reading speed on hardware that could do far more
The arithmetic finishes long before the next batch of numbers arrives. Speeding up the chip buys nothing when the wait is the problem, so the interesting question is how to widen the route rather than sharpen the cook.

2 · The obvious fix, and the wall it hits

Send the bits faster. Physics on a circuit board disagrees.

a clean signal, sent slowly the same signal, pushed hard every edge lands where it should smeared into something unreadable Long traces across a board have a ceiling on how fast you can drive them. So the speed dial is already close to the stop.
Bandwidth is how many wires there are multiplied by how fast each one runs, and the second factor runs into signal integrity on long board traces. With the speed dial near its limit, the only remaining term is the number of wires.

3 · The other dial

If you can’t run faster, run wider. Very much wider.

an ordinary memory channel 64 wires one stack of this kind 1024 wires Same modest speed per wire. Sixteen times as many of them. but a circuit board can only route a few hundred fast signals before it runs out of room
Widening the bus multiplies bandwidth without asking any single wire to go faster. The catch is that an ordinary board physically cannot route a thousand of them — which forces the next move.

4 · Where a thousand wires can actually fit

Build the floor out of silicon, and the memory has to move next door.

a slab of silicon, made the same way chips are the chip memory memory Silicon can carry thousands of microscopic traces. It just can’t carry them far. their reach is millimetres, not centimetres — so the memory has to ring the chip, and nowhere else
Fabricating the substrate with chip lithography is the only way to pack a thousand traces per stack. The price is distance: those traces reach millimetres, so the memory must sit right against the chip — which promptly runs the design out of floor space.

5 · Out of floor space

If you can’t grow outward, grow up.

wires drilled straight down through the dies eight or twelve dies high, in the footprint of one and each layer has to be bonded and tested — one bad die can spoil the whole stack heat from the bottom die also has to escape through every layer above it
Stacking gets tens of gigabytes into the small area the short traces allow, connected by vertical wires through the silicon. It buys density with manufacturing difficulty — yield and heat both get worse with height, and the manufacturers don’t publish the numbers.

6 · Keep this card

The whole thing on one index card.

the design = stack the memory dies, wired vertically + stand them on silicon, right beside the chip ∴ a bus that is very wide and unhurried Every part of it follows from one refusal: the wires would not go faster. and it only pays for work that streams enormous amounts of data through modest arithmetic
Picture to keep: not a faster pipe — a wider one. This memory doesn’t win the clock-speed race at all — it runs a thousand wires side by side, over a distance short enough that a thousand wires can physically exist.

Why it exists

You’ve watched an AI chatbot type its answer out at you, word by word, at a pace you can comfortably read along with. That’s odd when you think about it: the chip generating those words can do trillions of arithmetic operations per second, and a sentence is a few dozen words. It isn’t thinking that slowly. It’s waiting — for the model’s weights to arrive from memory, over and over, once per word.

Picture a busy restaurant where the chef can cook a dish in 5 seconds, but the pantry is in a building two blocks away. It doesn’t matter how fast the chef is — the kitchen runs at the speed of the runner fetching ingredients. To go faster, you don’t hire a faster chef. You move the pantry into the kitchen. High Bandwidth Memory — HBM — is the AI-chip version of that move. Instead of making the GPU compute faster, designers physically picked up the memory chips and stacked them millimeters away from the processor, on the same package. The chef and the pantry now share a counter. That short distance is a large part of why frontier models can be served at all.

Look at a die-shot of a GPU meant for AI — an H100, an MI300, a TPU package — and the visual is almost always the same. There’s a big square of compute logic in the middle, and right up against it, separated by a millimeter or two of silicon interposer, sit four to eight short rectangular towers. Those towers are the HBM stacks. They look out of place — like someone glued extra chips to the GPU — and that visual oddity is the whole story.

Regular computer memory doesn’t sit there. DDR sticks live a few centimeters away on the motherboard, connected by long copper traces. GDDR chips, used in gaming GPUs, sit a centimeter away on the same PCB. HBM sits next to the compute die on the same package, glued in by a special interposer, with thousands of wires running between them.

The reason is brutally simple: AI workloads don’t need more compute as much as they need more bandwidth — bytes per second from memory into the matrix multiplier — and the only way to get that many bytes that fast is to put the memory practically inside the chip.

Why it matters now

That chatbot typing at you is the everyday version of a much larger pattern: a lot of frontier-model serving is a memory-bandwidth problem dressed up as a compute problem. Generating one token for one user means reading the model’s weights and the KV cache out of memory and running them through the matrix multipliers exactly once. In that regime the multipliers are barely the bottleneck; the wires between memory and multipliers are. (Batch enough users together, or switch to training, and the balance shifts back toward compute — the bottleneck is a property of the workload, not of the chip.) This is why a chip’s HBM bandwidth — measured in terabytes per second, not gigabytes — ends up being one of the most-quoted numbers in any new AI silicon launch, sometimes ahead of the FLOP count.

It’s also why chip supply has gotten weird. The HBM market is dominated by a handful of DRAM manufacturers (publicly: SK hynix, Samsung, Micron), and constraints there now constrain who can build AI accelerators at all. The compute logic isn’t usually the gating part. The stacked memory is.

The short answer

HBM = stacked DRAM dies + through-silicon vias + silicon interposer + very wide, slow bus

Picture to keep: not a faster pipe — a wider one. HBM’s per-pin rate is unremarkable — comparable to a DDR5 stick, slower than GDDR; what it does is run a thousand wires side by side over a distance short enough that a thousand wires is physically possible.

HBM is just regular DRAM cells, rearranged. You take eight or twelve DRAM dies, stack them physically on top of each other, drill vertical wires straight through the silicon to connect them (those are TSVs), and then sit that whole tower on an interposer right next to the compute die. The bus to the compute die is deliberately slow per wire — HBM’s per-pin data rates are lower than GDDR’s, not higher — but the bus is very wide, often 1024 bits per stack. Bandwidth = width × rate per wire, and when you can’t push the rate further, you push the width.

How it works

Follow the chef-and-pantry problem through, one failed fix at a time.

Naive attempt: use normal memory. Bolt DDR sticks to the motherboard next to the accelerator, the way every server has done for decades. This fails immediately and quantitatively — the matrix units can consume bytes far faster than a handful of DDR channels can deliver them, so the chip spends most of its time idle. The runner is too slow, and the kitchen runs at the runner’s speed.

Fix 1: make the memory faster. Push the clock up.

Memory bandwidth for any DRAM technology is roughly bus width in bits × data rate per pin. Push the per-pin rate too hard and signal integrity falls apart on the long PCB traces — DDR5 sticks max out around 6–8 GT/s per pin in practice, GDDR pushes higher because the traces are shorter, but every step up the clock costs more power for diminishing returns. The physics is set by capacitance, inductance, and how loudly a wire couples to its neighbors. You can’t just print “10 GHz” on the box.

Fix 2: run wider instead. If you can’t run faster, run more wires in parallel. A DDR5 channel is 64 bits. A modern HBM stack is 1024 bits (split into 16 channels of 64 bits each). Instead of trying to win the GHz race, HBM wins the parallel-pins race. An H100-class part, depending on the SKU, carries five or six HBM stacks and moves a few terabytes per second. The same compute die paired with sticks of DDR would move a fraction of that.

But you can’t have 1024 wires per stack on a normal PCB.

This is where the interposer matters. A printed circuit board can route maybe a few hundred fast signals to a chip before you run out of layers and space. A silicon interposer — basically a thin extra slab of silicon underneath both the compute die and the HBM stacks — is fabricated with the same lithography used for chips, so it can carry thousands of microscopic traces packed tightly together. That’s the only way a 1024-bit-per-stack bus is physically buildable.

The interposer is also why HBM has to live so close. Those traces are microscopic, and their reach is short: push the distance and you start paying in signal integrity and power. Exactly how short depends on the signalling and packaging, and no clean threshold is published — but it’s millimetres, not centimetres, which is why HBM towers ring the compute die. They have nowhere else to go.

And now the memory has nowhere to live. Ringing the compute die with memory that must stay within a few millimetres leaves you very little floor space — but AI workloads want tens of gigabytes. That’s what forces the last move: if you can’t grow outward, grow up.

A single DRAM die only stores so many bits. To hit the tens of gigabytes per stack that AI workloads want, the dies are physically stacked — eight high, twelve high, sometimes more — and connected vertically through TSVs. Stacking trades cost and yield (you have to bond and test each layer; one bad die can ruin the stack) for density and very short vertical wires.

The standard account is that this stacking is genuinely hard manufacturing. Yield on a 12-high stack is worse than yield on a single die, and thermals are awkward — heat from the bottom die has to leave through layers of memory above it. The major manufacturers don’t publish yield numbers, so take “it’s hard” as the qualitative claim, not a specific figure.

Why this favors AI workloads specifically.

Go back to the chatbot. Generating one token for one user reads gigabytes of weights, runs each one through a multiply exactly once, and moves on. The arithmetic-to-memory ratio (the “arithmetic intensity”) is about as low as it gets — each weight is read and used once. That puts the workload on the memory-bandwidth side of the roofline, and HBM exists for exactly that regime.

The honest qualifier: not every transformer kernel lives there. Batch many users together, or train instead of serve, and the same weights get reused across many inputs, intensity climbs, and the kernel moves toward the compute ceiling. Training also drags in its own memory pressure (gradients, activations, optimizer state). So “bandwidth-bound” is a claim about a workload, not a law about transformers — but the workload that dominates interactive serving is squarely in HBM’s regime. And it’s still overkill for a laptop CPU running a database, where caches and a couple of DDR channels keep up fine.

The pantry-in-the-kitchen picture is right about distance and wrong about capacity: a real pantry gets bigger when you need more food, and HBM can’t. It’s stuck with whatever fits in a few stacks beside the die, which is why a model that doesn’t fit in HBM is a much bigger problem than a model that runs slowly — you fall off a cliff to system memory rather than sliding down a slope.

And that answers the die-shot: the towers are glued to the edge of the compute die because they have to be. A thousand-wire bus can only be built out of chip-scale traces, and chip-scale traces only reach a few millimetres. The layout isn’t a packaging choice — it’s the bus width made visible.

You started with HBM = stacked DRAM dies + TSVs + interposer + very wide, slow bus. What did this post add? — + "slow" is not a compromise, it's the design. Every part of HBM exists to make a wide bus physically possible, because the clock-speed lever was already pulled to the end. Stacking, TSVs, and the interposer aren’t three features; they’re three consequences of one decision to buy bandwidth by the wire instead of by the hertz.

Check yourself

Before you go — a vendor announces a new accelerator with 2× the FLOPs of the last one and the same HBM bandwidth. For a single-user LLM chat session, roughly how much faster should you expect token generation to be?

Answer

Barely faster, if at all. Generating one token for one user means streaming the entire weight set out of memory and multiplying it against a single vector — arithmetic intensity is low, so the kernel is sitting on the memory-bandwidth slope of the roofline, not under the compute ceiling. Doubling the ceiling doesn’t move a workload that never touches it. You’d expect a real speedup on training or on large-batch serving, where intensity is high enough to be compute-bound, and little on single-stream decoding. (Some gain can still leak in from better caches, fused kernels, or lower-precision formats that also cut bytes moved — but the FLOP number alone doesn’t buy it.) This is exactly why bandwidth numbers get quoted alongside, and sometimes ahead of, FLOP numbers in launch decks.

And one more — if wider buses are so much better than faster clocks, why doesn’t your laptop CPU use HBM?

Answer

Because the interposer is the expensive part and a CPU doesn’t need it. HBM’s width only pays off if your workload is memory-bandwidth bound; a CPU running a browser, a compiler, or a database has high enough cache hit rates and low enough streaming demand that a couple of DDR channels keep it fed. What you’d be buying instead is packaging cost, worse yield (a bad DRAM die can spoil a whole stack), harder thermals, and — critically — a memory capacity fixed at manufacture, since you can’t slot in more HBM the way you can add a DIMM. For general-purpose computing that’s a bad trade in every dimension.

Going deeper

One gap worth naming: there is no reliable public source for current HBM stack yields or per-stack manufacturing costs. The figures in circulation come from analyst firms (TrendForce, SemiAnalysis) rather than from the manufacturers, who don’t publish them.