Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why Spectre still isn't fully patched

Eight years after disclosure, new Spectre-class vulnerabilities keep landing. The reason isn't sloppy patching — the speculation being exploited is what makes modern CPUs fast, and the list of channels it leaks through has no end.

Security intermediate May 4, 2026 · updated Aug 25, 2026 · 14 min read

On this page

The picture version

Six pictures for a reader who has never thought about what a CPU does while it waits. The prose below fills in the seams the pictures skip.

1 · The problem

Every wall held. The memory leaked anyway.

a tab you didn’t choose someone else’s JavaScript, running in the sandbox the boundary memory it may not read the kernel, another origin’s page just read it? refused ask the CPU to guess — then time your own reads No permission check was ever violated. That is the whole difficulty.
An ordinary vulnerability breaks a rule the system meant to enforce. This one obeys every rule and reads the answer out of how long things took — which is why there is no single check to go and fix.

2 · Why the CPU guesses

It bins the wrong work. It doesn’t bin the smell.

a question the CPU can’t answer yet wait for it — and go idle guess, and run ahead guessed right: the work is already done guessed wrong: throw it away and “throw it away” turns out to mean two different things registers, program state rolled back completely nothing to see here cache lines, predictor entries left exactly as the guess left them and anyone can time a cache The right-hand box is not a bug. It is the speed.
Speculation is worth a large multiple in single-thread performance, so no vendor is giving it up. The undo is complete for everything the program can see, and absent for everything a stopwatch can see — that gap is the entire family of attacks.

3 · The walk

Teach it a habit, then break the habit once.

1 call it a thousand times with a legal index the branch predictor learns “this check always passes” 2 call it once with an illegal index the CPU assumes the check passes, reads the secret byte, and starts loading slot number “byte” 3 the check resolves, and everything is undone architecturally nothing happened — except one slot is now in the cache 4 time your own reads of all 256 slots one comes back fast, and its number is the byte fast slot 7 → the byte was 7 one byte per round, and the round is a loop
The secret never travels through a value the tab is allowed to keep. It travels as a choice of which slot to warm up, and the tab reads that choice back with a stopwatch.

4 · The three obvious fixes

Each one works. None of them finishes.

don’t let one domain’s guesses steer another’s flush or partition the predictors at every boundary crossing covers the predictor paths it names, and no others cost: syscalls, context switches undo the cache too, not just the registers the theoretically right answer, and it exists — in simulation what ships instead is per-variant: silicon fixes for named leaks and the cache is one channel of many stop guessing altogether the kernel does exactly this on a few chosen lines of code as a general policy it would undo decades of performance so it stays a scalpel, not a switch Every shipped fix patches a channel. None of them patches the mechanism. which is why the answer to “is it patched?” is always “which one?”
The real defence is a patchwork — predictor barriers, rewritten indirect calls, kernel page-table isolation, blunter browser clocks, scheduling that keeps distrusting tenants off the same core. Each entry closes a path someone demonstrated, which is a different thing from closing the class.

5 · Why it never ends

The variants are cells in a grid.

ways to make it speculate ×  places the residue shows up branch predictor return stack store-to-load forwarding gather instructions cache TLB store buf fill buf ports ✓✓ ✓ ✓ ✓ ??? ???? ???? ???? ✓  a demonstrated path, and a mitigation shipped for it ?  a cell that is still only a potential paper Cells get patched. The grid is a property of the design.
The rows and columns here are illustrative, not a census — the real point is the shape. Every named variant is one cell someone got to first, so mitigating it shrinks the grid by one and leaves the rest live.

6 · Keep this card

Two halves, and neither one comes out.

Spectre = speculative execution + a side channel that survives the rollback the speculation can’t go it is most of your single-thread performance, so nobody is volunteering to give it back the channels can’t be listed once and closed “shared microarchitectural state” is a category that gains new members every year A real fix would be a different CPU, not a patch.
So “is Spectre patched?” has no yes-or-no answer. Named variants are patched; the property that generates them is the design. The useful question is the threat-model one instead — is anything untrusted running on this hardware at all?

Why it exists

You have a browser tab open that you didn’t choose — an ad frame, an embedded widget, something running JavaScript that someone else wrote. It’s sandboxed: it can’t read your files, can’t touch other tabs, can’t call the kernel. Every boundary the browser promises is intact. And yet, starting in January 2018, that tab could in principle read memory it had no permission to read, without violating a single one of those boundaries — by asking the CPU to guess, and then timing how long its own memory accesses took.

That tab is the running example for this post. Hold onto it.

The consequences are the reason a security update can leave the same machine measurably slower than it was the week before: since January 2018 every mainstream OS, browser, and hypervisor has been piling on mitigations for a family of CPU bugs that started with two papers named Spectre and Meltdown. The patches cost performance — sometimes a few percent, sometimes much more for system-call-heavy workloads — and they keep coming. Inception, Downfall, Reptar and GhostRace landed across 2023 and 2024; in 2025 Training Solo re-opened cross-domain Spectre v2 attacks on a long list of Intel parts, and VMScape leaked from a guest virtual machine into its hypervisor on every AMD Zen generation up to and including Zen 5.

The interesting question isn’t what Spectre is. It’s why the industry can’t just fix it and move on the way it does with an ordinary CVE.

The short version: Spectre doesn’t exploit a bug in the usual sense — a typo in a memcpy, a missing bounds check, a parser that trusts its input. It exploits the intended behavior of essentially every high-performance CPU built since the late 1990s. Modern processors are fast largely because they refuse to wait: they guess what comes next, run it, and throw the work away if the guess was wrong. That guessing — speculative execution — is worth a large multiple in single-thread performance. Spectre is the discovery that “throw the work away” is incomplete: the architectural state is reverted, but microarchitectural side effects — which cache lines got loaded, which predictor entries got updated — persist, and an attacker can read them through timing.

You can’t remove the bug, because the bug is the optimization. You can only narrow the channels through which it leaks, one variant at a time, and pay for it in throughput.

Why it matters now

Spectre-class issues are why cloud providers behave strangely about SMT / hyperthreading on shared-tenant hardware — some disable it for certain workloads, others restrict cross-tenant pairings, because two threads sharing a core’s branch predictors and caches is exactly the topology these attacks want. They’re also why browser JavaScript engines coarsened their high-resolution timers and switched SharedArrayBuffer off at the start of 2018 — and the terms of its return say more than the switch-off did. Shared memory is available again, but only to a page that is cross-origin isolated: one that has used HTTP headers to declare that it embeds no cross-origin content which hasn’t opted in — which is also what lets the browser put it in a process of its own. The capability wasn’t restored, it was re-priced. Throughout, the threat model assumes attacker code is already running on your machine — in a tab, in a Lambda function, in a JS sandbox — and is only trying to read memory it shouldn’t reach.

Most importantly, the meta-story keeps repeating. Researchers find a new way to coerce the CPU into speculating across a security boundary; vendors ship microcode or a compiler flag; months later someone finds the next variant. Each year’s crop isn’t new physics — it’s new paths through the same physics. Until CPUs are designed with speculation isolated by security domain, the trickle is structural rather than incidental.

The short answer

Spectre = speculative execution + a microarchitectural side channel that survives the rollback

Picture to keep: a line cook who starts the dish before the order is confirmed. Wrong order? He bins the food — but the kitchen still smells of what he cooked, and someone standing outside the door can tell what it was. Where the picture breaks: a smell is vague, while the CPU’s residue is precise and addressable — the attacker doesn’t sniff, they time 256 specific memory locations and read off exactly one byte.

The CPU runs ahead of the program, executing instructions it isn’t yet sure are needed. Guessed wrong, it discards registers and pipeline state — but the cache, the branch predictor, and other shared buffers keep the fingerprints of what it touched. An attacker tricks the CPU into speculating through a memory access it shouldn’t make, then reads those fingerprints through timing. The “fix” is to plug channels one at a time, because the speculation itself is too valuable to give up.

How it works

The attack, concretely

The original Spectre variant (Kocher et al., 2018, arXiv:1801.01203) is the cleanest illustration. Imagine code in a more privileged domain that takes an integer index x from your tab and does:

if (x < array1_size) {
    y = array2[array1[x] * 4096];
}

The bounds check looks airtight. An out-of-range x returns immediately.

But the branch predictor doesn’t know the law — it learns from history. Your tab calls this repeatedly with valid x, training the predictor to expect “branch taken.” Then it calls once with an out-of-range x. The CPU predicts taken, speculatively dereferences array1[x] — a read of memory your tab has no right to — uses the byte it found as an index into array2, and begins loading the corresponding cache line. Then the bounds check resolves, speculation is squashed, registers are restored. Architecturally, nothing happened.

But the cache line is still loaded. Your tab now times its own accesses to each page of array2. One comes back fast — the one that got speculatively touched — and its index is the secret byte. Repeat, one byte at a time.

That’s Spectre v1: bounds-check bypass via the branch predictor. Note what the tab never did: it never violated a permission check. It asked a question it was allowed to ask, and measured how long the answer took.

Why the obvious fixes don’t close it

Walk the three things you’d naturally try.

Naive fix 1: don’t speculate across security boundaries. If the problem is that predictor state trained in one domain steers speculation in another, partition or flush that state at the boundary.

Why it’s only partial: this is what Intel’s IBRS, IBPB, and STIBP microcode features do, and each has a measurable cost on syscalls, context switches, and SMT pairs — you’re deliberately throwing away the learned history that made prediction work. Compiler-side, retpoline rewrites indirect branches into a pattern the predictor can’t influence, at a cost per call. Both work for the predictor paths they cover. Neither is a statement about speculation in general.

Naive fix 2: revert the cache state too. The rollback already restores registers; make it restore microarchitectural state as well, and the channel closes at the source.

Why it’s only partial: this is the theoretically right answer, and academic designs exist — SafeSpec and InvisiSpec both do it, evaluated in simulation. What vendors actually ship is narrower: silicon fixes for named variants and microcode barriers for named predictors, not a cache that forgets. You’d have to track which lines were loaded speculatively, undo them on every squash, and do it inside the cycle budget that made speculation worth having. And the cache is only one channel — the TLB, store buffers, fill buffers, and port contention are all still there. The disclosures that keep landing on current parts are the evidence that nobody has closed this at the source: VMScape reaches cores designed years after Spectre was public.

Naive fix 3: turn speculation off. It works. The kernel does exactly this on selected paths, using speculation barriers or index masking after sensitive bounds checks.

Why it’s only partial: as a general policy it’s unaffordable. Modern CPU performance largely is the speculation; disabling it broadly would undo decades of microarchitecture.

So the real portfolio is a patchwork: KPTI for Meltdown, retpoline for indirect branches, microcode predictor barriers, reduced browser timer precision (which is what took away your tab’s measuring instrument), hypervisor core scheduling so VMs don’t share SMT siblings, and per-variant fixes as researchers find new channels.

Why new variants keep landing

The original paper named two variants and explicitly anticipated more. Researchers promptly started cataloguing the speculation primitives a CPU offers (indirect branches, return stacks, store-to-load forwarding, gather instructions, transactional aborts) and the side channels available (L1/L2/L3 cache, TLB, store buffers, line-fill buffers, port contention). Cross every primitive with every channel and you get a matrix; each cell is a potential paper. Foreshadow, MDS/RIDL/Fallout/ZombieLoad, LVI, Retbleed, Downfall (gather-data sampling), Inception (AMD return-stack injection), Reptar (Intel’s redundant-prefix issue), GhostRace (speculative race conditions), Training Solo, VMScape — these are mostly cells in that matrix, not new physics.

The pattern: someone proves this speculative path leaks through that buffer; vendors mitigate that specific path; the matrix shrinks slightly and the rest stays live. That’s why “is Spectre patched?” is the wrong question. Individual cells are patched. The matrix is a property of the design.

What about hardware-fixed CPUs?

Vendors have been quietly redesigning, one named variant at a time. Intel started folding fixes into silicon with the Whiskey Lake and Cascade Lake parts of 2018–19 — Meltdown and L1TF handled in hardware rather than microcode — and each later generation has added a few more hardware controls and enumeration bits saying which attacks no longer apply to this part. AMD’s Zen cores were never vulnerable to Meltdown, and grew variants of their own instead: Inception, which AMD mitigates with microcode on Zen 3 and Zen 4, and VMScape, whose AMD advisory lists Zen 1 through Zen 5 as potentially affected. ARM’s cores have a similarly mixed record. The trend is “more of the known leaks closed in silicon,” not “a new architecture without speculation” — nobody is volunteering to give up that much single-thread performance, and the closer you look the more the channels multiply. Speculation interacts with caches, predictors, prefetchers, memory ordering, and SMT in ways that offer no single chokepoint to defend.

Show the seams

You started with Spectre = speculative execution + a side channel that survives the rollback. What did this post add that explains the eight years? — + the two halves are separately unremovable. Speculation can’t go, because it’s most of your performance. And the side channels can’t be enumerated once and closed, because “shared microarchitectural state” is a category with new members every year. A fix would have to make speculation leave no trace at all — which is a different CPU, not a patch.

Check yourself

Before you go — a colleague proposes that since Spectre needs precise timing, adding random jitter to the browser’s clock makes the attack impossible. Is “impossible” the right word?

Answer

No — “more expensive” is. Noise doesn’t remove the signal, it lowers the signal-to-noise ratio, and an attacker can recover the signal by repeating the measurement and averaging. Spectre reads one byte at a time anyway, so it’s already a loop; making each byte take a thousand trials instead of ten is a real cost but not a wall.

That’s why the browser response was broader than jitter: reduce timer resolution and remove SharedArrayBuffer (which let attackers build their own high-resolution clock from a counter thread, and which came back only for cross-origin-isolated pages) and move cross-origin content into separate processes via site isolation. Notice the last one is the only structural move — it doesn’t degrade the attacker’s instrument, it removes the secret from the address space they can reach. When you see a mitigation that makes an attack noisier, ask what it would cost the attacker to just try more times.

And: your workload is a single-tenant server running only code you wrote, and mitigations cost you 15%. Is disabling them reckless?

Answer

Not automatically — it’s a threat-model question, and this is one of the few places where the honest answer is “it depends, and here’s on what.”

Spectre needs an attacker executing code on your machine. If nothing untrusted runs there, the main precondition is absent. But check the assumptions carefully, because “only my code” is stronger than most people mean: do you run a JIT or any interpreter over user-supplied input (a query language, a template engine, a regex over hostile strings)? Do you unmarshal untrusted data into a parser with attacker-influenced control flow? Does anything on the box handle secrets belonging to different users of your service, such that a gadget in your own code could leak across them? Any yes puts attacker-shaped computation on your hardware.

The general lesson matters more than the verdict: Spectre mitigations aren’t a hygiene checkbox to be maximized, they’re a purchase, and what you’re buying is isolation between mutually distrusting things sharing a core. If nothing on the machine distrusts anything else, you’re paying for a boundary you don’t have.

Going deeper