Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why CPU clock speeds stopped climbing

Around the mid-2000s, GHz numbers on CPUs flatlined while core counts started growing. The reason isn't engineering laziness — it's that switching a transistor costs energy, and energy turns into heat you can't get rid of fast enough.

Science intro Apr 29, 2026 · updated Aug 25, 2026 · 9 min read

On this page

The picture version

Four pictures for a reader who remembers the megahertz race. The prose below fills in the seams the pictures skip.

1 · The problem

The number on the box climbed for twenty years, then stopped.

clock speed mid-2000s 100 MHz → 3 GHz and then flat, for twenty years — still single-digit GHz today core count, meanwhile Intel didn’t run out of ideas. Transistors kept shrinking. something else hit a limit — and the industry quietly switched from “go faster” to “go wider”
Curves are the shape of the history, not plotted data. The quantity that ran out is power density — watts per square millimetre of die — and the limit it creates is the power wall.

2 · Where the watts come from

An energy per flip, times how often you flip.

P ≈ α · C · V² · f C — what you have to drive V — supply voltage f — clock rate α — fraction switching C·V² is an energy per flip. Multiply by a rate and you get power. so doubling the clock doubles the power — at the same voltage and underneath it all, leakage: modern transistors trickle current even when nothing is switching. multiply by a few billion and you get watts.
Note the square. That exponent is about to be the whole story — first working for the industry, then against it.

3 · The deal that broke

For decades, shrinking let voltage fall too. Then it didn’t.

while voltage could fall power per mm², flat more transistors, faster clocks, same heat once it hit its floor power per mm², climbing the same square now works against you Dennard scaling broke before Moore’s Law did. They were separate promises. the standard account: voltage couldn’t keep dropping without noise swamping the signal and leakage exploding — a gradient, not a cliff, somewhere in the mid-2000s
Moore’s Law is a count of transistors. Dennard scaling was the promise those transistors would fit in the same thermal envelope. The second one expired first, and everything since is a response to that.

4 · Keep this card

The whole thing on one index card.

the power wall = the point where switching power + leakage exceeds what you can carry away as heat ∴ the V² is the part that ended the megahertz race
Picture to keep: a hotplate the size of a postage stamp. Every extra gigahertz is another burner turned up on it, and the only way heat leaves is through a stack of metal and moving air. Where the picture is wrong: a hotplate’s output is roughly proportional to the dial, and a CPU’s isn’t — voltage has to rise to sustain a higher clock, and power goes as the square, so the last 20% of clock speed costs far more than the first.

Why it exists

You know the sound: you open one more browser tab or start a video export, and the laptop fan winds up like a hairdryer while the underside gets uncomfortably warm. Then, a minute later, the work is still going but the fan settles and the machine feels slower. That’s not the software giving up. That’s the chip deliberately slowing itself down because it can’t get rid of heat fast enough. The same effect, at industry scale, is why CPU clock speeds stopped climbing twenty years ago.

If you remember computers from the 1990s, you remember the megahertz race. Every new CPU bragged louder about clock speed. 100 MHz, 500 MHz, 1 GHz, 2 GHz, 3 GHz — the number on the box went up roughly in step with calendar time, and faster clock meant a faster machine in a way you could feel.

Then the curve broke. Somewhere around the mid-2000s, mainstream desktop CPUs stalled in the 3–4 GHz range and basically stayed there. Two decades later, a top-end consumer chip in 2026 still boosts to single-digit gigahertz. Meanwhile core counts went from one, to two, to four, to dozens. The industry quietly switched from “go faster” to “go wider.”

The interesting question isn’t what happened — it’s why physics forced it. Intel didn’t run out of ideas. Transistors kept getting smaller and cheaper. Something else hit a wall. The quantity is power density — watts per square millimetre of die — and the limit it creates has a name: the power wall.

Why it matters now

Every modern conversation about compute scaling — datacenter siting, GPU thermal design, training-run economics, even why your laptop fan kicks on during an inference call — is downstream of this same physics. GPUs are wide rather than fast, and the power ceiling is a large part of why — alongside the fact that their workloads parallelise well enough to make width pay. TPUs are wide. Apple’s M-series chips are wide. And a large part of why an AI training cluster needs a substation rather than a wall outlet is that once single-thread headroom ran out, performance started being bought by the kilowatt instead of by the gigahertz.

Software engineers feel this every time they write code that doesn’t parallelize. Single-thread performance still improves — branch prediction, cache, wider issue, better compilers — but the days of waiting two years for your serial loop to magically run twice as fast are over.

The short answer

power wall = the point where switching power + leakage exceeds what you can carry away as heat

Picture to keep: a hotplate the size of a postage stamp. Every extra gigahertz is another burner turned up on it, and the only way heat leaves is through a stack of metal and moving air. The clock speed you can keep is whatever the hotplate can shed without cooking itself — the same reason your laptop fan is the thing that decides how fast your export finishes.

A digital CPU burns power in two main ways. Switching — every time a transistor flips it charges or discharges a tiny capacitance, costing on the order of C·V² of energy per flip. Do that f times a second, across a fraction α of the transistors, and the power is roughly α·C·V²·f. Leakage — even when nothing is switching, modern transistors are leaky enough that current trickles through them all the time. Multiply by a few billion transistors and you get watts. Watts turn into heat. Heat has to leave the chip, or the chip melts. The clock speed you can sustain is whatever lets the heat budget balance.

How it works

The naive attempt: just turn the clock up. This is what the industry did happily for thirty years, and it’s still what your intuition wants. The transistors will switch faster; nothing stops you. Follow the chain of what breaks, in order, and you get the whole story.

1. Switching energy scales with frequency.

Take the basic dynamic-power equation that every chip designer carries in their head: P ≈ α · C · V² · f. C is the capacitance you have to drive (set by the wires and gates), V the supply voltage, f the clock frequency, and α the fraction of transistors actually switching in a given cycle. Note what each side is: C·V² is an energy per flip, and multiplying by a rate turns it into power. Crucially, doubling the clock doubles the power at the same voltage — because you’re now paying that switching cost twice as often per second.

For a long time, this was fine, because of Dennard scaling: as transistors shrank, you could drop the voltage in step. P drops with V², so even though you packed more transistors and ran them faster, total power per square millimeter stayed roughly flat. This was the deal that powered the megahertz race.

2. Dennard scaling broke before Moore’s Law did.

Moore’s Law is a count: transistors per chip, doubling on a regular cadence. Dennard scaling is a power claim: those transistors fit in the same thermal envelope. The two rode together for decades, but they’re separate promises, and Dennard’s broke first — somewhere around the 90 nm / 65 nm process nodes in the early-to-mid 2000s.

The standard account is that voltage couldn’t keep dropping without making transistors unreliable: thermal noise becomes comparable to the signal, leakage explodes, and threshold voltages have a physical floor. Once V stops dropping but transistor counts keep doubling, the C·V²·f budget for the whole chip blows up. You’re suddenly trying to dissipate hundreds of watts from a thumbnail-sized die.

I’m being deliberately hand-wavy about the exact node where this “broke” because the answer is gradient, not a cliff edge — but the consensus is that by the mid-2000s the old game was over.

3. Heat removal is the actual ceiling.

You could, in principle, just clock the chip faster anyway. The transistors will switch. The problem is that the heat has to go somewhere. A modern CPU die is roughly the size of a postage stamp. The path the heat takes to leave is: silicon → integrated heat spreader → thermal paste → heatsink baseplate → fins → moving air (or water). Each interface has a thermal resistance, and resistance times power equals temperature drop.

Push the junction temperature high enough and the part stops being reliable — vendors publish a maximum in the region of 100 °C, and in ordinary operation a chip that reaches it throttles rather than destroys itself. Either way the ceiling binds: how many watts can you shove through that stack of materials before the clock has to come down? For a desktop air cooler that’s on the order of a couple of hundred watts sustained. For a datacenter GPU on a custom liquid loop, it’s more, but still bounded.

This is why you can’t fix the power wall by being clever with the chip alone. You can build a faster engine, but you can’t build a heatsink that breaks the second law of thermodynamics.

Why “go wide” works around it.

A core running at 3 GHz instead of 6 GHz uses much less than half the power, because you can also drop the voltage when the clock is lower, and power scales with V². Two cores at 3 GHz can do about as much arithmetic per second as one (hypothetical) core at 6 GHz, but use far less power to do it — if the workload parallelizes. So the industry pivoted: more cores, wider SIMD, more specialized accelerators, all at modest clocks. GPUs are the limit case — thousands of slow lanes instead of a few fast ones.

The hotplate picture is right about the ceiling and wrong about one thing: a hotplate’s heat output is roughly proportional to how hard you turn the dial, and a CPU’s isn’t. Because voltage has to rise to sustain a higher clock, and power goes as V², pushing the clock up is super-linear in heat. That’s why the wall arrives so abruptly — the last 20% of clock speed costs far more than the first 20%.

The whole “AI needs absurd amounts of electricity” story is a downstream consequence. Once parallelism is your only lever, you scale by adding more silicon and more cooling, and the bill is paid in megawatts.

So — that fan winding up on your laptop is the wall, in miniature. You started with power wall = the point where switching power + leakage exceeds what you can carry away. What did this post add? — + the V² is the part that ended the megahertz race. As long as shrinking transistors let voltage fall alongside the shrink (Dennard scaling), the square worked for you. Once voltage hit its floor, the same square started working against you, and “go wide” became the lever that was left.

Going deeper