Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why matrix multiplication is the bottleneck of modern ML

Modern ML is mostly one operation in a trench coat. Understanding why matmul dominates explains hardware, software, and why GPUs eat the world.

Math intermediate Apr 29, 2026 · updated Aug 25, 2026 · 10 min read

On this page

The picture version

Five pictures for a reader who has never wondered what a graphics card is for. The prose below fills in the seams the pictures skip.

1 · The problem

One operation is doing almost all of the work, everywhere.

writing the next wordpainting a pictureunlocking with your face multiplying big tables of numbers together A graphics card is, essentially, a factory built for this one job. which raises an obvious question: why did everything converge on the same operation?
Very different-looking capabilities bottom out in the same calculation, repeated at enormous scale. That convergence is not a coincidence, and it isn’t only about the mathematics — the hardware had a vote.

2 · The expectation that turned out to be wrong

“Design what’s right; the hardware will catch up.”

how fast it can calculate how fast data can reach it time → Arithmetic got cheap far faster than delivering the numbers to it did.
For a long stretch it really was true that anything you invented got faster on its own. What broke that was the two curves separating: calculating became enormously cheaper while fetching the numbers did not keep pace.

3 · What that gap actually selects for

Operations that use each number once are starved. This one doesn’t.

add two lists together fetch two numbers, do one addition, throw them away the machines sit waiting on the belt multiply two tables fetch a block once, use every value in it many times the belt keeps the whole floor busy The work grows faster than the data does — which is exactly what a starved machine wants. doubling the size of the tables multiplies the arithmetic eightfold but the data only fourfold
An operation that does one sum per number fetched can never outrun the belt feeding it. Multiplying tables is the shape where one fetched block earns many multiplications, so the arithmetic units stay fed — provided the work is arranged to reuse each block while it is close at hand.

4 · Then the loop closes

Shapes that run fast get built on. Chips get built for the shapes people build.

this shape runs fast today so more people build with it so chips specialise for it so it runs faster still so what survives isn’t only what works best in principle — it is what the machine rewards
An idea that happens to run well attracts more work, which justifies hardware tuned for it, which makes it faster again. The filter isn’t purely mathematical — and this is the part that genuinely is a design history, because chips are designed by people responding to what researchers are already doing.

5 · Keep this card

The whole thing on one index card.

a modern model ≈ a stack of table multiplications + something cheap in between the expensive half is the half the hardware is built for Not because nothing else could work — because this is what the machine feeds well. which is worth remembering whenever someone says an architecture “won”
Picture to keep: a factory floor of thousands of identical multiply-and-add machines, with one conveyor belt feeding them. This is the job shaped so that a single trip down the belt keeps every machine busy; most other operations leave much of the floor idle waiting on the belt.

Why it exists

Every time ChatGPT writes a word, Midjourney paints a pixel, or your phone unlocks with Face ID, the hidden grunt work behind the scenes is the same one operation: multiplying enormous tables of numbers together. That’s matrix multiplication — “matmul.” It’s why a single high-end GPU costs more than the rest of a gaming PC combined: the GPU is essentially a factory built to do this one calculation millions of times in parallel. Take this operation away and modern AI doesn’t just slow down — it disappears.

Open the profiler on almost any modern model — a transformer, a diffusion model, a ResNet from a decade ago — and you’ll see the same picture. One operation dominates the FLOP count, most of the memory traffic, and most of the wall-clock time: matrix multiplication, multiplying two 2D tables of numbers so that each output entry is a dot product of a row and a column. Everything else — the activations, the normalizations, the softmaxes — is a small share of the arithmetic budget.

That’s a strange thing to discover. Why would a field as varied as “vision, language, audio, robotics, protein folding” all collapse onto a single linear-algebra primitive? It’s not because researchers picked it; it’s because everything else stopped scaling, and matmul didn’t.

Why it matters now

This one fact silently shapes decisions you’re already making:

The short answer

modern neural net ≈ a stack of (matmul + cheap nonlinearity)

Picture to keep: a factory floor of thousands of identical multiply-and-add machines, and one conveyor belt feeding them. Matmul is the job shaped so that a single trip down the belt keeps every machine busy; most other operations leave much of the floor idle waiting on the belt. (Where the picture breaks: the belt isn’t the only way to waste the floor — a matmul that’s too small, or one tiled badly, leaves the machines idle too. High intensity is a property of the kernel, not automatically of the operation.)

A neural network layer, stripped of branding, is: take a vector of inputs, multiply it by a learned weight matrix, add a bias, apply a cheap elementwise function (ReLU, GELU). Stack a hundred of those. That’s it. Even attention is matmul wearing a different hat — attention(Q, K, V) = softmax(Q · Kᵀ / √d) · V is two matmuls with a softmax in between, and producing Q, K, and V in the first place is three more. (Softmax itself isn’t elementwise — each output depends on the whole row — but it’s cheap next to the multiplies around it.)

So when you train or run a model, you are mostly asking your hardware to multiply matrices. Faster matmul = faster everything.

How it works

The naive expectation. Hardware gets faster every year, so whatever operation a researcher invents will get faster too. Design the model you think is right; the silicon will catch up. That was roughly true through the 1990s, and it’s the intuition most people still carry.

Why it breaks. Compute got cheap much faster than memory got fast. That gap is the memory wall, and it means the number that decides an operation’s fate isn’t how many FLOPs it does — it’s how many FLOPs it does per byte it drags out of memory. Adding two vectors does one FLOP for two loads and a store: the chip finishes instantly and then sits idle waiting on memory. Feed a modern accelerator that workload and you use a rounding error’s worth of its arithmetic capability. Every operation that isn’t dense enough to hide its own memory traffic quietly fell off the Pareto frontier — not because it was slower in theory, but because the machine spent its time waiting.

What survives the filter. Multiplying two N×N matrices is O(N³) arithmetic against O(N²) data. Written well — tiled so that a block of the matrix is loaded once and reused across many multiplies — every value pulled from HBM earns its trip many times over. That ratio, arithmetic intensity, grows with N: matmul is the rare operation that gets better at hiding the memory wall the bigger you make it. That’s the survivor — and note the “written well,” because a badly tiled matmul falls right back onto the memory wall like everything else.

Then the loop tightens. Around the late 2000s and early 2010s, deep learning beat hand-engineered features at vision, then speech, then language, and the winning models all shared a shape: many layers, each layer a big matmul. Convnets (a convolution can be unrolled into a matmul via im2col, and for years that was the common implementation), RNNs, transformers — different architectures, same primitive. Then GPUs added dedicated matmul units (NVIDIA shipped tensor cores in Volta, 2017), and the gap widened again: a tensor-core matmul can be several times faster than the same FLOPs run through generic vector lanes. So a model built out of matmuls runs far faster than the same FLOP budget spent on anything else, and researchers followed the speed. Hardware selected for matmul, then architectures selected for the hardware.

The seam. Matmul’s dominance isn’t a law of nature — it’s a feedback loop between hardware and architectures. There are operations (sparse attention, structured matrices, state-space models) that are mathematically cheaper but currently slower in practice because the hardware isn’t shaped for them. Whether that loop ever breaks — whether something dethrones matmul — is genuinely open. Mamba and friends are the most credible recent challengers, but as of early 2026, transformers still win at the frontier, and the frontier still runs on matmul.

One more thing worth naming: the theoretical exponent of matrix multiplication is below 3 — Strassen’s algorithm is O(N^2.807), and there’s a long line of asymptotically faster algorithms going down toward ~2.37. Almost none of them show up in production ML kernels. The usual explanation — worse constants, worse numerical stability, worse cache and tiling behaviour at the sizes we actually run — is the standard account rather than something I can point at a benchmark for. What’s not in doubt is what ships: production GEMM is the classical cubic algorithm, tiled and hand-tuned to within an inch of its life. Which is the whole thesis in one sentence: the asymptotically cheaper algorithm loses to the one shaped like the machine.

So: back to the opening. Why does the grunt work behind ChatGPT, Midjourney, and Face ID collapse onto the same operation, when the three problems have nothing in common? Not because anyone chose it. Because the memory wall filters hard, and matmul is the shape that comes through it.

You started with modern neural net ≈ a stack of (matmul + cheap nonlinearity). What did this post add? — + and the memory wall, not elegance, is what selected it. Matmul didn’t win on FLOP count; it won on FLOPs-per-byte, and every “why is my model slow” question is really “which of my operations lost that ratio.”

Check yourself

Before you go — a colleague replaces a big dense layer with a sparse one that does 10× fewer FLOPs, and reports the model got slower on the same GPU. What happened?

Answer

Cutting FLOPs cuts the numerator of arithmetic intensity without cutting memory traffic proportionally — unstructured sparsity still drags data (weights plus index metadata) across the bus, it just multiplies less of it. The kernel slides from compute-bound toward memory-bound, so the GPU spends its time waiting on HBM instead of multiplying, and most of the 10× FLOP saving evaporates. On top of that, an unstructured sparse kernel typically can’t use the tensor cores, so the FLOPs it does perform run on the slower general-purpose lanes. Structured sparsity that the hardware explicitly supports is a different story — which is exactly the point: “mathematically cheaper” and “faster on this hardware” are separate claims, and only the second one is about the machine you own.

And one more — a 70B model generating one token at a time for a single user gets nowhere near the GPU’s advertised FLOP number, but the same GPU serving 64 users at once gets much closer. Why would batching change the arithmetic at all?

Answer

It doesn’t change the arithmetic — it changes the ratio. With one user, each weight matrix is loaded from HBM and multiplied against a single vector, so you get one pass of arithmetic per byte loaded: intensity is terrible and the chip is memory-bound. Batch 64 users and you load the same weights once but multiply them against 64 vectors — a matrix-matrix product instead of a matrix-vector one. Bytes moved stay roughly flat, FLOPs go up 64×, arithmetic intensity goes up with them, and the kernel crosses over into compute-bound territory where the advertised number lives. It’s the same reason big-N matmul beats small-N matmul, applied to serving.

Going deeper