Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why does speculative decoding exist?

A small fast model guesses, a big slow model checks. Somehow you get the big model's exact output, faster. The trick isn't cleverness — it's that your GPU was already sitting idle.

AI & ML intermediate Apr 29, 2026 · updated Aug 25, 2026 · 14 min read

On this page

The picture version

Six pictures for a reader who has never wondered why a chatbot types at reading speed. The prose below fills in the seams the pictures skip.

1 · The problem

That word-by-word pace isn’t thinking. It’s a 140 GB read, per word.

The capital of France is Paris. 140 GB140 GB140 GB140 GB140 GB140 GB six words, six complete sweeps through every weight in the model meanwhile, the arithmetic units mostly sitting there with nothing to do You paid for the arithmetic. The reading is what you’re waiting on. the hardware is asking for more work per byte fetched — this trick takes it up on that
Generating one word means streaming the model’s entire weights out of memory, and then doing it again for the next word. The bottleneck is reading, not calculating — which leaves a great deal of the machine idle.

2 · The obvious fix, and why it isn’t one

Just use the small model. Fine here — and you can’t tell when it isn’t.

the easy sentence both models say “Paris” the small one is 70× cheaper to run, and right the sentence you actually cared about they disagree, quietly and nothing in the output tells you so You didn’t want a cheap answer. You wanted the expensive model’s answer. A speed-up you can’t audit is a downgrade. so keep the small model — but never let it have the final say
A small model handles easy text just as well and costs a fraction to run. The problem is that you cannot tell from the outside which case you are in — so the small model needs a supervisor, not a promotion.

3 · The asymmetry everything rests on

Checking four guesses costs the big model about what checking one costs.

scoring one word read all the weights a little arithmetic scoring four at once read all the weights — the same once four times a little The expensive half didn’t change. Only the cheap half got bigger. true while reading dominates, and while the number of guesses stays small
Once the weights have been pulled out of memory, using them against four candidate positions instead of one costs almost nothing extra. That asymmetry is the entire opening — and it disappears once the machine is busy enough to be limited by arithmetic instead.

4 · The move

Let the small model run ahead, then check the whole run in one glance.

the small model guesses ahead capitalofFranceisLyon one pass of the big model scores every one of them agreedagreedagreedagreed no keep this run stop here, and write “Paris” itself One expensive pass produced five words instead of one. and the correction is free: the same pass already worked out what the right word was
The cheap model proposes a run of words; the expensive one scores them all in a single pass, keeps the stretch it agrees with, and writes the first word it doesn’t. Every pass now yields somewhere between one word and the whole run — and one rejection ends the run, which is why guessing ever further ahead pays less and less.

5 · The part that makes it honest

“Accept it if they match” would quietly change the answer. So the rule is cleverer.

the tempting rule “keep the guess whenever the big model would have said it too” quietly tilts the output toward whatever the small model likes the rule that actually works keep the guess with a probability set by how much the two models disagree and when you reject, draw the replacement from exactly what the big model was missing the result is the big model’s own output, provably
Accepting whenever the two models happen to agree would reweight the output toward the small model — the exact quality leak the whole design exists to avoid. The published rule accepts with a computed probability and, on rejection, samples from the difference between the two, so the sequence you get is the one the big model would have produced alone.

6 · Keep this card

The whole thing on one index card.

the trick = a cheap model guesses several words ahead + the expensive one checks the run in one pass + an accept rule that leaves the answer unchanged you are not buying a cheaper answer — you are buying the same one, sooner Take away either half and the trick stops being one. without the rule you just have a worse model; without idle hardware, two models to pay for
Picture to keep: a junior drafting a sentence on a whiteboard while the expert reads over their shoulder — the expert checks the whole line at a glance, crosses out from the first mistake onward, and writes the next word themselves. Where it breaks: a real expert would be swayed by the draft, and this one provably isn’t. On a busy server the idle capacity is already spoken for, and the whole trade stops paying.

Why it exists

You ask a chatbot a question and watch the answer arrive one word at a time, at roughly reading speed. It feels like the model is thinking, deliberating, choosing. It isn’t. That pace is a hardware fact, and once you see what’s producing it, the pace looks less like thought and more like waste.

Here’s our running example, and we’ll keep it for the whole post: a 70B-parameter LLM generating the sentence “The capital of France is Paris.”

Serve that model densely, one conversation at a time — the plain case, no batching tricks — and generating one token of that sentence forces the GPU to read every weight of the model out of memory. A 70B-parameter model in 16-bit precision is ~140 GB of weights. To produce the next token, the GPU reads all 140 GB again. And again. Six tokens, six full sweeps through 140 GB. The arithmetic involved per token is small; the reading is what kills you. This is the memory-bandwidth wall, and it’s why your single-stream inference is slow even on a card that brags about petaflops.

Here’s the part that should really bother you: while all that reading is happening, the GPU’s compute units are mostly idle. You paid for petaflops. On dense, low-batch decode you’re using a small fraction of them. The hardware is begging you to give it more arithmetic to do per byte read.

Speculative decoding is the move that takes the hardware up on that offer.

You probably assume the catch is quality: a small model helps out, so you must be getting a slightly cheaper, slightly worse answer. It’s the opposite. The output distribution is provably identical to running the big model alone — the small model never gets a vote it can’t be overruled on. The thing you’re trading away isn’t accuracy; it’s idle silicon.

Why it matters now

The major open inference stacks ship it as a documented feature: vLLM, TensorRT-LLM, and llama.cpp all have it. Whether a given hosted API uses it is mostly not disclosed, so I won’t claim a list. For agent workloads — long, decode-heavy chains of tool calls and reasoning — it’s one of the few ways to cut wall-clock latency without dropping to a smaller model and eating the quality loss.

It also matters because it composes (mostly — real stacks have feature-by-feature gaps) with the other tricks. MoE, KV-cache reuse, prompt caching — speculative decoding stacks on top of them.

The short answer

speculative decoding = small draft model + big model verifies in parallel

Picture to keep: a junior drafting a sentence on a whiteboard while the expert reads over their shoulder — the expert checks the whole line in one glance instead of dictating word by word, crosses out from the first mistake onward, and writes the next word themselves. The expert’s standards never drop; they just stopped doing the typing.

A cheap “draft” model guesses the next K tokens. The big “target” model then runs one forward pass that scores all K guesses simultaneously, accepts the longest prefix it agrees with, and produces one bonus token of its own. If the draft is right most of the time, you get several tokens per expensive forward pass instead of one — and the output distribution is provably the same as what the big model would have produced alone.

How it works

The mechanism rests on one asymmetry that most people miss the first time they hear about this technique: in the memory-bound regime, and for small K, scoring K tokens in parallel costs the big model roughly what scoring one costs.

That’s because, in the memory-bound regime, the cost of one decode step is dominated by streaming the weights through the memory hierarchy, not by the matmuls themselves. Once a weight tile is loaded into the GPU’s caches it can be reused against a length-1 input or a length-K input at roughly the same wall-clock cost, until K gets large enough that you finally become compute-bound. So the big model can verify “did I agree with all K of these guesses?” in basically one forward pass.

Now build the technique by breaking things. Keep our sentence in mind: the 70B model is producing “The capital of France is Paris.”

Naive attempt: just use the small model. A 1B sibling gets this sentence right — it’s easy — and it’s dramatically cheaper to run, since it has ~70× fewer weights to stream per token. Ship that.

Why it breaks: you didn’t want a 1B answer, you wanted a 70B answer. On this sentence the two agree; on the hard sentence you actually care about they won’t, and from the outside you cannot tell which case you’re in. A speedup you can’t audit isn’t a speedup, it’s a downgrade.

Fix 1: let the big model check the small model’s work. The 1B model drafts K tokens ahead — say capital of France is — and the 70B model runs one forward pass over the prompt plus those four drafted tokens. That single pass hands you the target’s probability distribution at every one of those positions at once, because of the asymmetry above. Now you can ask, position by position, “would I have said that?”

Why that breaks: the draft will eventually be wrong. Suppose it drafts capital of France is Lyon. You can’t accept the whole block, and throwing the whole block away wastes the three tokens that were fine.

Fix 2: accept the longest prefix, then correct. Walk left to right and take the run of tokens the target agrees with; stop at the first disagreement. And you get one extra token essentially free at the stopping point, out of the same forward pass — if all K were accepted, you sample from the target’s distribution at position K+1, which the verify pass already computed; if you rejected at position j, you sample the replacement token from the target’s (corrected) distribution at j. Either way the big model’s pass produced somewhere between 1 and K+1 tokens instead of exactly 1. Then loop.

Why that breaks — the subtle one: “agrees with” is easy to define for greedy decoding (accept the draft token iff it’s also the target’s argmax) but not for sampling. If you sample at temperature and accept whenever the draft’s token happens to match, you have quietly reweighted the distribution toward whatever the small model likes, and your output is no longer the big model’s output. This is exactly the quality leak you were trying to avoid.

Fix 3: a rejection-sampling rule that’s provably unbiased. From the Leviathan et al. paper: accept the drafted token with probability min(1, p_target(x) / p_draft(x)), and on rejection sample from the residual (p_target − p_draft)+, normalized. The joint distribution over generated sequences is provably the target’s.

That last bit is the part that surprises engineers, and it’s the payoff for the whole chain. You’re not approximating the big model. You’re not trading quality for speed. The output distribution is identical to what the target would have produced alone. (Implementations can still differ by tiny amounts due to floating-point numerics — the guarantee is at the math level, not at the bit level.) The draft model is just a guess source; the target retains full veto power.

This is where the over-the-shoulder analogy earns its keep — and where it breaks. A real expert reading a junior’s draft is influenced by it; anchoring is a real thing. The target model isn’t. The accept/reject rule is constructed so that the draft can change how fast you get the answer and never which answer you get. That’s the difference between a helpful colleague and a provably-unbiased one.

Where the speedup comes from

Two ways to think about it, both useful:

The two failure modes are symmetric. If the draft is too bad, acceptance rates collapse, and you’re paying for the small model’s forward passes plus the big model’s, with little to show. If the draft is too good (e.g. a 7B drafting for a 70B), the small model itself is now slow, and the cost of running it eats into the savings. Picking a draft is a Goldilocks problem.

A real-world acceptance rate of 60–80% on natural prose is typical with a well-matched draft, giving end-to-end speedups of roughly 2–3× on memory-bound inference. Reported numbers vary widely by workload — code, structured output, and chat all behave differently — and there is no canonical public benchmark to point at, so treat “2–3×” as a rough order of magnitude rather than a guarantee.

Where it stops working

Speculative decoding helps memory-bound inference. If you’re already compute-bound — large batch sizes serving many users in parallel, or very small models where the weights aren’t the bottleneck — there are no idle flops to recover, and the math stops being free. This is why hosted providers sometimes turn it on for low-batch / low-traffic regimes and off when traffic is heavy. The exact crossover point depends on the model, the hardware, and the batch size, and no clean public number exists for where modern inference servers flip the switch.

You started with speculative decoding = small draft model + big model verifies in parallel. What did the failure chain add? — + an unbiased accept/reject rule, and a hardware regime where it pays. The accept/reject rule is what turns “a cheaper approximation” into “the same answer, sooner”; the memory-bound regime is what makes the verification pass nearly free. Take away either one and the trick stops being a trick: without the rule you’ve just got a worse model, and without idle flops you’ve just got two models to pay for.

Which answers the thing we opened with. That word-by-word pace never was deliberation — it was a memory bus being read end to end, once per word. Speculative decoding doesn’t make the model think faster. It just stops making it re-read the whole library to write each one.

Check yourself

Before you go — you swap your draft model for a bigger, better one. Acceptance rate climbs from 70% to 90%. Is the system necessarily faster? What would you actually compare?

Answer

Not necessarily, and this is the Goldilocks problem from above. Higher acceptance means more accepted tokens per expensive target pass, which is pure win on the target side. But a bigger draft model costs more per drafted token, and you pay that cost K times per cycle whether the tokens get accepted or not. The comparison that matters is total wall-clock per output token: (cost of K draft steps + cost of 1 target pass) / (expected tokens accepted per cycle). If your draft went from 1B to 7B against a 70B target, you multiplied the draft-side cost roughly sevenfold to buy maybe 1.3× more tokens per cycle — likely a loss. The extreme case makes it obvious: a draft model identical to the target has 100% acceptance and zero speedup.

And one more — your provider’s docs say speculative decoding is enabled, but you see no latency improvement at all during business hours and a clear one at 3am. Nothing about your prompt changed. What’s the most likely explanation?

Answer

You’re not in the memory-bound regime during the day. Speculative decoding recovers idle compute, and at high concurrency the server is already filling those flops with other users’ requests via batching — the GPU is closer to compute-bound, so there’s nothing to recover and verifying K positions genuinely costs more than verifying one. At 3am your request is nearly alone on the card, decode is bandwidth-starved, and the free capacity is there to exploit. This is also why the technique sits awkwardly next to continuous batching: both are after the same idle silicon, and on a busy server the batching is already claiming it.

Going deeper