Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why GPUs ended up running AI even though they were built for graphics

GPUs were designed to shade pixels. Then the same hardware turned out to be exactly what neural networks needed. That isn't luck — graphics and deep learning make the same demand of silicon: identical arithmetic, millions of times over, with nothing to branch on.

Science intro Apr 29, 2026 · updated Aug 25, 2026 · 11 min read

On this page

The picture version

Six pictures for a reader who has never thought about what a chip is made of. The prose below fills in the seams the pictures skip.

1 · The odd fact

The part you wanted for gaming is the part a datacenter wanted for AI.

one graphics card drawing explosions at sixty frames a second predicting the next word in a sentence Two jobs with nothing in common — except to the silicon, where they’re one job. the rest of this explains what the silicon sees that we don’t
Rendering a scene and running a neural net look like unrelated problems from the outside. At the level the hardware cares about they have the same shape, which is why one part ended up doing both.

2 · One chef, or a thousand line cooks

A CPU spends its silicon being clever. A GPU spends it on arithmetic.

CPU a few strong cores wrapped in caches, prediction, reordering brilliant at one unpredictable job GPU thousands of weak lanes almost no control logic each hopeless at that — unbeatable at a thousand identical ones A GPU is not a fast computer. It is a very wide one.
Almost all of a CPU’s silicon goes to making a single unpredictable instruction stream fast — caches, branch prediction, reordering. A GPU deletes nearly all of that and spends the transistors on multipliers instead, which is only affordable because of the constraint in the next scene.

3 · The bargain

The lanes are cheap because they all have to run the same step.

one instruction, issued once, for every lane multiply & add different numbers in each lane — same step, same moment Ask for a branch and the deal breaks: the lanes that take it and the lanes that don’t run in turn — the branch executes twice, half the chip idle each time which is why fast GPU code avoids data-dependent control flow rather than merely disliking it
Thousands of lanes are affordable only because they share one instruction stream, so the chip skips almost all per-lane control logic. That’s the whole bargain, and also the bill: a workload with real branching pays for the lanes and can’t use them.

4 · Why the fit is so good

Shading a million pixels and multiplying two big matrices are the same job.

a frame each pixel: same little program nobody waits on their neighbour a matrix multiply each output: same little dot product nobody waits on their neighbour Graphics needed this shape first. Deep learning arrived and found the seat warm. CUDA in 2007 opened the door; AlexNet in 2012 walked through it on two consumer gaming cards
A pixel shader and a matmul output element are both “one identical, branch-free computation, times millions, with no dependencies between them.” The hardware that wins at that shape is the same hardware either way — and it had already been built and mass-produced for gamers.

5 · The part people miss

The moat isn’t the arithmetic. It’s the loading dock.

a pantry down the hall memory cores a narrow path, hidden behind big caches a loading dock memory, sitting right against it lanes terabytes per second, straight in Training is mostly limited by moving weights, not by multiplying them. graphics was bandwidth-hungry too — it streams textures and framebuffers constantly — so the memory system built to feed pixel shaders now feeds matrix multiply units
Thousands of idle lanes are worthless, so the real constraint is how fast weights and activations reach them. GPUs inherited an extremely wide memory system from graphics, which had the same appetite for exactly the same reason.

6 · Keep this card

The whole thing on one index card.

GPU = thousands of slow arithmetic lanes + wide memory to keep them fed + everyone runs the same instruction That third clause is the bargain: it buys the lanes, and it bills any branchy workload. custom AI chips aren’t trying to be better GPUs — they keep this shape and drop the graphics legacy
Picture to keep: a CPU is one brilliant chef who can improvise; a GPU is a thousand line cooks who can only follow the same recipe step at the same moment, standing next to a loading dock instead of a pantry. Terrible for a bespoke dinner. Unbeatable when the order is a thousand identical plates.

Why it exists

If you tried to buy a graphics card in the last few years, you ran into something odd: you were bidding against AI companies. The part you wanted for playing games was the same part a datacenter wanted for training neural networks, and the datacenter had more money. That’s a strange coincidence on its face — rendering explosions and predicting the next word in a sentence have nothing to do with each other. Except the hardware doesn’t think so.

If you’d told a graphics engineer in 1999 that the chip they were designing to draw triangles in Quake would, twenty-five years later, be the most expensive and most fought-over component in a datacenter, they would have laughed. GPUs were not built for science. They were built so that the same work could be done, independently, on every pixel of a screen at sixty frames per second. (In 1999 that work was still mostly fixed-function; the fully programmable per-pixel shader arrived a couple of years later. The shape — one operation, applied to millions of independent pixels — was there from the start.)

And yet today essentially every frontier neural network — every LLM, every diffusion model, and most large recommenders — is trained on GPUs or on accelerators built from the same throughput-first playbook, like Google’s TPUs. The interesting question isn’t that this happened, it’s why the fit is so absurdly good. Graphics and deep learning look like completely different problems. Why does the same hardware win at both?

The short version: both problems are, deep down, the same shape — do the same arithmetic to a giant pile of numbers, all at once, with no branching. The hardware that solved one already solved the other. We just didn’t notice for a while.

Why it matters now

If your product calls a model, a GPU is somewhere in your cost structure. The reason an inference call costs what it costs, the reason fine-tuning a 70B model is expensive, the reason a startup’s runway is partly a function of NVIDIA’s gross margin — it comes back to the fact that the chips that do dense linear algebra at scale and at reasonable cost are either descendants of pixel shaders or purpose-built machines designed on the same throughput-over-latency principle.

Understanding why graphics hardware became AI hardware also tells you what an alternative would have to look like. Custom AI silicon (TPUs, Trainium, MTIA, Cerebras wafers) isn’t trying to “be a better GPU” — it’s trying to keep the parts of the GPU that matter for neural nets and drop the parts that exist purely because of legacy graphics.

The short answer

GPU = thousands of slow arithmetic lanes + wide memory + "everyone runs the same instruction" execution model

Picture to keep: a CPU is one brilliant chef who can improvise; a GPU is a thousand line cooks who can only follow the same recipe step at the same moment, standing next to a loading dock instead of a pantry. Terrible for a bespoke dinner. Unbeatable when the order is a thousand identical plates.

A GPU is not a fast computer. A single GPU “core” is much weaker than a CPU core — slower clock, dumber branch prediction, smaller caches per lane. What a GPU has is many of them, all forced to run the same instruction at the same time, fed by an unusually fat pipe to memory. That’s exactly the recipe for shading a million pixels. It’s also exactly the recipe for multiplying two big matrices, which is what a neural network mostly is.

How it works

Naive attempt: train the network on a CPU. It’s the general-purpose computer; that’s what general-purpose is for. And it works — it’s what everyone did before roughly 2010. It’s just hopelessly slow, and the reason it’s slow is instructive, because it isn’t that the CPU’s arithmetic is weak.

Why it breaks. A CPU spends most of its silicon on being clever about one instruction stream: branch prediction, out-of-order execution, deep caches, speculative loads. All of that exists to make unpredictable, branchy, one-thing-at-a-time code fast. A neural network’s forward pass is the opposite workload — the same multiply-accumulate, repeated over millions of independent numbers, with no branches to predict and no dependencies to reorder. Every transistor the CPU spent on cleverness is wasted, and the handful it spent on actual multipliers is the only part doing work.

The fix already existed, built for something else. Graphics had this exact problem thirty years earlier, and the hardware that solved it was sitting in gaming PCs the whole time. Three threads make the fit clear.

1. Graphics is embarrassingly parallel, and its heaviest stages are matrix math.

To render a 3D scene, the GPU does roughly: take a list of vertices, multiply each one by a 4×4 transformation matrix to get screen coordinates; then for every pixel covered, run a small program (a shader) to decide its color. The work on one pixel doesn’t depend on the work on the next pixel. There are millions of them. They all run the same program. This is the textbook definition of data parallelism.

Underneath, the arithmetic-heavy stages are small matrix multiplies and dot products — transforming vertices, lighting calculations — while others, like texture sampling and color blending, are fixed-function work that is parallel in the same way without being matmul. Late-1990s GPUs did this with a handful of fixed-function pipelines; as the decade turned and shaders became programmable, vendors kept widening that arrangement until a modern part has thousands of arithmetic lanes running in parallel, every frame, forever. The architecture that fell out of this is called SIMT: many threads, all running the same instruction in lockstep on different data.

2. Neural networks turn out to be the same shape.

A forward pass through a transformer is, to a first approximation, a stack of large matrix multiplications interleaved with cheap element-wise operations (additions, nonlinearities, normalizations). Backprop is more matrix multiplications. Training a model for months is, mostly, doing matmul forever. (See why matmul is the bottleneck.)

Matmul is about as embarrassingly parallel as numerical computing gets: every output element is an independent dot product. There are no branches, no sequential dependencies inside the multiply, no need for elaborate control flow. It’s the exact workload SIMT was built for — except instead of “one instruction over a million pixels,” it’s “one instruction over a million matrix tiles.”

The realization that GPUs could run general numerical code crystallized around 2007 with NVIDIA’s CUDA, which exposed the GPU as a programmable parallel machine instead of a fixed-function pixel pipeline. Deep learning’s takeoff moment — the AlexNet ImageNet result in 2012 — ran on two consumer GeForce GTX 580 cards and CUDA. That wasn’t a coincidence; it was among the first widely visible cases of “the graphics chip is also the math chip.”

3. Wide memory, not fast clocks, is the actual moat.

Modern training is mostly limited by how fast you can move weights and activations between memory and the arithmetic units, not by how fast the units themselves can multiply. (See memory bandwidth.) Data-center GPUs ship with HBM — DRAM stacks bonded next to the chip — delivering terabytes per second of bandwidth. That bandwidth exists because graphics, too, is bandwidth-bound: you’re streaming textures and framebuffers constantly. The same memory subsystem that fed pixel shaders now feeds matrix multiply units.

The CPU world chose latency: small, fast caches, complex prediction, lots of silicon spent on making one thread fast. The GPU world chose throughput: many slow lanes, very wide memory, no per-thread cleverness. AI happens to want throughput, not latency.

What’s vestigial, and where the seams show.

Not everything on a GPU is useful for AI. On a graphics-first part, texture units, raster operators, ray-tracing cores and video encoders sit mostly idle during a training run, however useful they are for the workloads they were built for. Data-center parts have already shed much of that: NVIDIA describes H100 as largely not a graphics chip, with only a small fraction of it graphics-capable. NVIDIA has gradually paved over this by adding Tensor Cores, which are essentially “matmul accelerators bolted into the same SIMT envelope.” A modern training GPU is mostly a matmul engine wearing a graphics chip’s skin. Custom AI chips like TPUs go further and drop the graphics legacy entirely: a TPU is, roughly, a giant systolic array with a tiny bit of control logic around it.

One gap worth naming. The exact die-area split between “graphics legacy” and “AI-relevant” silicon on, say, an H100 or a B200 isn’t publicly broken down, and the boundary is fuzzy anyway because units like memory controllers and schedulers serve both. The standard account — that AI workloads are now the design target and graphics is increasingly the side gig — is well-supported by NVIDIA’s own roadmap statements, but no precise breakdown exists to quote.

The line-cook picture is right about the parallelism and wrong about one thing: the cooks aren’t independent. They move in lockstep, and if the recipe says “if the sauce is too thin, do X,” the cooks who need X and the cooks who don’t take turns — the branch runs twice, once with half the kitchen idle. That’s SIMT divergence, and it’s why GPU code avoids data-dependent control flow rather than merely disliking it.

The deeper reason this all worked is older than either field: when a workload is “the same arithmetic, repeated, with no dependencies,” the hardware that wins is the hardware that gives up everything else for parallel arithmetic and bandwidth. Graphics demanded that shape first. Deep learning showed up later and discovered the seat was already warm.

So — that graphics card you couldn’t buy. You weren’t bidding against datacenters only because of a shortage or a coincidence of manufacturing. You were bidding against them because the thing you wanted it to do and the thing they wanted it to do are, at the level the silicon cares about, the same job.

You started with GPU = thousands of slow arithmetic lanes + wide memory + "everyone runs the same instruction". What did this post add? — + that third clause is the whole bargain. The lockstep requirement is what makes thousands of lanes affordable, because the chip can skip nearly all per-lane control logic and spend the transistors on arithmetic instead. It’s also the bill: any workload with real branching pays for those lanes and can’t use them.

Going deeper