Why Spectre still isn't fully patched
Eight years after disclosure, new Spectre-class vulnerabilities keep landing. The reason isn't sloppy patching — the speculation being exploited is what makes modern CPUs fast, and the list of channels it leaks through has no end.
On this page
The picture version
Six pictures for a reader who has never thought about what a CPU does while it waits. The prose below fills in the seams the pictures skip.
1 · The problem
Every wall held. The memory leaked anyway.
2 · Why the CPU guesses
It bins the wrong work. It doesn’t bin the smell.
3 · The walk
Teach it a habit, then break the habit once.
4 · The three obvious fixes
Each one works. None of them finishes.
5 · Why it never ends
The variants are cells in a grid.
6 · Keep this card
Two halves, and neither one comes out.
Why it exists
You have a browser tab open that you didn’t choose — an ad frame, an embedded widget, something running JavaScript that someone else wrote. It’s sandboxed: it can’t read your files, can’t touch other tabs, can’t call the kernel. Every boundary the browser promises is intact. And yet, starting in January 2018, that tab could in principle read memory it had no permission to read, without violating a single one of those boundaries — by asking the CPU to guess, and then timing how long its own memory accesses took.
That tab is the running example for this post. Hold onto it.
The consequences are the reason a security update can leave the same machine measurably slower than it was the week before: since January 2018 every mainstream OS, browser, and hypervisor has been piling on mitigations for a family of CPU bugs that started with two papers named Spectre and Meltdown. The patches cost performance — sometimes a few percent, sometimes much more for system-call-heavy workloads — and they keep coming. Inception, Downfall, Reptar and GhostRace landed across 2023 and 2024; in 2025 Training Solo re-opened cross-domain Spectre v2 attacks on a long list of Intel parts, and VMScape leaked from a guest virtual machine into its hypervisor on every AMD Zen generation up to and including Zen 5.
The interesting question isn’t what Spectre is. It’s why the industry can’t just fix it and move on the way it does with an ordinary CVE.
The short version: Spectre doesn’t exploit a bug in the usual sense — a typo in a memcpy, a missing bounds check, a parser that trusts its input. It exploits the intended behavior of essentially every high-performance CPU built since the late 1990s. Modern processors are fast largely because they refuse to wait: they guess what comes next, run it, and throw the work away if the guess was wrong. That guessing — speculative execution — is worth a large multiple in single-thread performance. Spectre is the discovery that “throw the work away” is incomplete: the architectural state is reverted, but microarchitectural side effects — which cache lines got loaded, which predictor entries got updated — persist, and an attacker can read them through timing.
You can’t remove the bug, because the bug is the optimization. You can only narrow the channels through which it leaks, one variant at a time, and pay for it in throughput.
Why it matters now
Spectre-class issues are why cloud providers behave strangely about
SMT / hyperthreading
on shared-tenant hardware — some disable it for certain workloads, others
restrict cross-tenant pairings, because two threads sharing a core’s branch
predictors and caches is exactly the topology these attacks want. They’re also
why browser JavaScript engines coarsened their high-resolution timers and
switched SharedArrayBuffer off at the start of 2018 — and the terms of its
return say more than the switch-off did. Shared memory is available again, but
only to a page that is
cross-origin isolated:
one that has used HTTP headers to declare that it embeds no cross-origin
content which hasn’t opted in — which is also what lets the browser put it in
a process of its own. The capability wasn’t restored, it was re-priced. Throughout, the threat
model assumes attacker code is already running on your machine — in a tab,
in a Lambda function, in a JS sandbox — and is only trying to read memory it
shouldn’t reach.
Most importantly, the meta-story keeps repeating. Researchers find a new way to coerce the CPU into speculating across a security boundary; vendors ship microcode or a compiler flag; months later someone finds the next variant. Each year’s crop isn’t new physics — it’s new paths through the same physics. Until CPUs are designed with speculation isolated by security domain, the trickle is structural rather than incidental.
The short answer
Spectre = speculative execution + a microarchitectural side channel that survives the rollback
Picture to keep: a line cook who starts the dish before the order is confirmed. Wrong order? He bins the food — but the kitchen still smells of what he cooked, and someone standing outside the door can tell what it was. Where the picture breaks: a smell is vague, while the CPU’s residue is precise and addressable — the attacker doesn’t sniff, they time 256 specific memory locations and read off exactly one byte.
The CPU runs ahead of the program, executing instructions it isn’t yet sure are needed. Guessed wrong, it discards registers and pipeline state — but the cache, the branch predictor, and other shared buffers keep the fingerprints of what it touched. An attacker tricks the CPU into speculating through a memory access it shouldn’t make, then reads those fingerprints through timing. The “fix” is to plug channels one at a time, because the speculation itself is too valuable to give up.
How it works
The attack, concretely
The original Spectre variant (Kocher et al., 2018,
arXiv:1801.01203) is the cleanest
illustration. Imagine code in a more privileged domain that takes an integer
index x from your tab and does:
if (x < array1_size) {
y = array2[array1[x] * 4096];
}
The bounds check looks airtight. An out-of-range x returns immediately.
But the branch predictor
doesn’t know the law — it learns from history. Your tab calls this repeatedly
with valid x, training the predictor to expect “branch taken.” Then it
calls once with an out-of-range x. The CPU predicts taken, speculatively
dereferences array1[x] — a read of memory your tab has no right to — uses
the byte it found as an index into array2, and begins loading the
corresponding cache line. Then the bounds check resolves, speculation is
squashed, registers are restored. Architecturally, nothing happened.
But the cache line is still loaded. Your tab now times its own accesses to
each page of array2. One comes back fast — the one that got speculatively
touched — and its index is the secret byte. Repeat, one byte at a time.
That’s Spectre v1: bounds-check bypass via the branch predictor. Note what the tab never did: it never violated a permission check. It asked a question it was allowed to ask, and measured how long the answer took.
Why the obvious fixes don’t close it
Walk the three things you’d naturally try.
Naive fix 1: don’t speculate across security boundaries. If the problem is that predictor state trained in one domain steers speculation in another, partition or flush that state at the boundary.
Why it’s only partial: this is what Intel’s IBRS, IBPB, and STIBP microcode features do, and each has a measurable cost on syscalls, context switches, and SMT pairs — you’re deliberately throwing away the learned history that made prediction work. Compiler-side, retpoline rewrites indirect branches into a pattern the predictor can’t influence, at a cost per call. Both work for the predictor paths they cover. Neither is a statement about speculation in general.
Naive fix 2: revert the cache state too. The rollback already restores registers; make it restore microarchitectural state as well, and the channel closes at the source.
Why it’s only partial: this is the theoretically right answer, and academic designs exist — SafeSpec and InvisiSpec both do it, evaluated in simulation. What vendors actually ship is narrower: silicon fixes for named variants and microcode barriers for named predictors, not a cache that forgets. You’d have to track which lines were loaded speculatively, undo them on every squash, and do it inside the cycle budget that made speculation worth having. And the cache is only one channel — the TLB, store buffers, fill buffers, and port contention are all still there. The disclosures that keep landing on current parts are the evidence that nobody has closed this at the source: VMScape reaches cores designed years after Spectre was public.
Naive fix 3: turn speculation off. It works. The kernel does exactly this on selected paths, using speculation barriers or index masking after sensitive bounds checks.
Why it’s only partial: as a general policy it’s unaffordable. Modern CPU performance largely is the speculation; disabling it broadly would undo decades of microarchitecture.
So the real portfolio is a patchwork: KPTI for Meltdown, retpoline for indirect branches, microcode predictor barriers, reduced browser timer precision (which is what took away your tab’s measuring instrument), hypervisor core scheduling so VMs don’t share SMT siblings, and per-variant fixes as researchers find new channels.
Why new variants keep landing
The original paper named two variants and explicitly anticipated more. Researchers promptly started cataloguing the speculation primitives a CPU offers (indirect branches, return stacks, store-to-load forwarding, gather instructions, transactional aborts) and the side channels available (L1/L2/L3 cache, TLB, store buffers, line-fill buffers, port contention). Cross every primitive with every channel and you get a matrix; each cell is a potential paper. Foreshadow, MDS/RIDL/Fallout/ZombieLoad, LVI, Retbleed, Downfall (gather-data sampling), Inception (AMD return-stack injection), Reptar (Intel’s redundant-prefix issue), GhostRace (speculative race conditions), Training Solo, VMScape — these are mostly cells in that matrix, not new physics.
The pattern: someone proves this speculative path leaks through that buffer; vendors mitigate that specific path; the matrix shrinks slightly and the rest stays live. That’s why “is Spectre patched?” is the wrong question. Individual cells are patched. The matrix is a property of the design.
What about hardware-fixed CPUs?
Vendors have been quietly redesigning, one named variant at a time. Intel started folding fixes into silicon with the Whiskey Lake and Cascade Lake parts of 2018–19 — Meltdown and L1TF handled in hardware rather than microcode — and each later generation has added a few more hardware controls and enumeration bits saying which attacks no longer apply to this part. AMD’s Zen cores were never vulnerable to Meltdown, and grew variants of their own instead: Inception, which AMD mitigates with microcode on Zen 3 and Zen 4, and VMScape, whose AMD advisory lists Zen 1 through Zen 5 as potentially affected. ARM’s cores have a similarly mixed record. The trend is “more of the known leaks closed in silicon,” not “a new architecture without speculation” — nobody is volunteering to give up that much single-thread performance, and the closer you look the more the channels multiply. Speculation interacts with caches, predictors, prefetchers, memory ordering, and SMT in ways that offer no single chokepoint to defend.
Show the seams
- There is no single performance number, and be wary of anyone who quotes one. Published figures vary wildly — single-digit percent for compute-bound work, 20%+ for syscall-heavy workloads on early patches, less on newer hardware with silicon mitigations. The right number depends on CPU generation, kernel version, which mitigations are enabled, and the workload. If you need a number for a real decision, benchmark the actual machine; don’t quote this post.
- The shape transfers across vendors; the specifics don’t. Intel, AMD, ARM, IBM POWER, and Apple Silicon each have their own mitigation stack and their own vulnerability matrix, and a mitigation named in one vendor’s advisory usually means nothing on another’s part. What holds everywhere is the shape — speculation is the mechanism, the channel is microarchitectural state, the fix is per-variant. For a specific machine, the vendor’s own advisory for that part is the only thing that answers the question.
- “Unpatched” doesn’t mean “exploitable in your threat model.” These attacks need attacker-controlled code running on your hardware, decent timing, and often a specific gadget in the victim. That’s a real threat for browsers and multi-tenant cloud, and much less pressing for a single-tenant box running only your code.
You started with Spectre = speculative execution + a side channel that survives the rollback. What did this post add that explains the eight years?
— + the two halves are separately unremovable. Speculation can’t go, because
it’s most of your performance. And the side channels can’t be enumerated once
and closed, because “shared microarchitectural state” is a category with new
members every year. A fix would have to make speculation leave no trace at
all — which is a different CPU, not a patch.
Check yourself
Before you go — a colleague proposes that since Spectre needs precise timing, adding random jitter to the browser’s clock makes the attack impossible. Is “impossible” the right word?
Answer
No — “more expensive” is. Noise doesn’t remove the signal, it lowers the signal-to-noise ratio, and an attacker can recover the signal by repeating the measurement and averaging. Spectre reads one byte at a time anyway, so it’s already a loop; making each byte take a thousand trials instead of ten is a real cost but not a wall.
That’s why the browser response was broader than jitter: reduce timer
resolution and remove SharedArrayBuffer (which let attackers build their
own high-resolution clock from a counter thread, and which came back only for
cross-origin-isolated pages) and move cross-origin
content into separate processes via site isolation. Notice the last one is
the only structural move — it doesn’t degrade the attacker’s instrument, it
removes the secret from the address space they can reach. When you see a
mitigation that makes an attack noisier, ask what it would cost the attacker
to just try more times.
And: your workload is a single-tenant server running only code you wrote, and mitigations cost you 15%. Is disabling them reckless?
Answer
Not automatically — it’s a threat-model question, and this is one of the few places where the honest answer is “it depends, and here’s on what.”
Spectre needs an attacker executing code on your machine. If nothing untrusted runs there, the main precondition is absent. But check the assumptions carefully, because “only my code” is stronger than most people mean: do you run a JIT or any interpreter over user-supplied input (a query language, a template engine, a regex over hostile strings)? Do you unmarshal untrusted data into a parser with attacker-influenced control flow? Does anything on the box handle secrets belonging to different users of your service, such that a gadget in your own code could leak across them? Any yes puts attacker-shaped computation on your hardware.
The general lesson matters more than the verdict: Spectre mitigations aren’t a hygiene checkbox to be maximized, they’re a purchase, and what you’re buying is isolation between mutually distrusting things sharing a core. If nothing on the machine distrusts anything else, you’re paying for a boundary you don’t have.
Famous related terms
- Meltdown —
Meltdown ≈ Spectre's cousin that read past the privilege check during speculation— affected Intel and some other designs (notably not AMD), and was cleaner to fix: Kernel Page-Table Isolation removes the kernel mapping from the user page table. Disclosed alongside Spectre in January 2018. - KPTI —
KPTI = separate page tables for user and kernel + a more expensive context switch— the canonical Meltdown mitigation, and the source of most of the syscall-cost story. - Retpoline —
retpoline = indirect call rewritten as a return-trampoline the predictor can't poison— Google’s compiler-side mitigation for Spectre v2. - MDS —
MDS = the same trick + internal buffers instead of the cache— the family covering RIDL, Fallout, and ZombieLoad. - Why CPUs have cache levels —
cache hierarchy = small-and-fast + large-and-slow— the timing difference the whole attack is measured with. - ASLR —
ASLR = randomize the address layout + force the attacker to leak it first— Spectre-class leaks hand over that first step, which is one reason ASLR’s cost-raising story has weakened.
Going deeper
- Kocher et al., Spectre Attacks: Exploiting Speculative Execution (2018) — for the question “what is the precise primitive?”; the variant zoo since is mostly elaboration on this.
- Hill et al., On the Spectre and Meltdown Processor Security Vulnerabilities — for the question “what does this mean for how processors should be designed?”, written as a retrospective by computer architects rather than security researchers.
- Lipp et al., Meltdown: Reading Kernel Memory from User Space (2018) — the rabbit hole: for the question “why was the twin bug killable when this one wasn’t?”