Why does speculative decoding exist?
A small fast model guesses, a big slow model checks. Somehow you get the big model's exact output, faster. The trick isn't cleverness — it's that your GPU was already sitting idle.
On this page
The picture version
Six pictures for a reader who has never wondered why a chatbot types at reading speed. The prose below fills in the seams the pictures skip.
1 · The problem
That word-by-word pace isn’t thinking. It’s a 140 GB read, per word.
2 · The obvious fix, and why it isn’t one
Just use the small model. Fine here — and you can’t tell when it isn’t.
3 · The asymmetry everything rests on
Checking four guesses costs the big model about what checking one costs.
4 · The move
Let the small model run ahead, then check the whole run in one glance.
5 · The part that makes it honest
“Accept it if they match” would quietly change the answer. So the rule is cleverer.
6 · Keep this card
The whole thing on one index card.
Why it exists
You ask a chatbot a question and watch the answer arrive one word at a time, at roughly reading speed. It feels like the model is thinking, deliberating, choosing. It isn’t. That pace is a hardware fact, and once you see what’s producing it, the pace looks less like thought and more like waste.
Here’s our running example, and we’ll keep it for the whole post: a 70B-parameter LLM generating the sentence “The capital of France is Paris.”
Serve that model densely, one conversation at a time — the plain case, no batching tricks — and generating one token of that sentence forces the GPU to read every weight of the model out of memory. A 70B-parameter model in 16-bit precision is ~140 GB of weights. To produce the next token, the GPU reads all 140 GB again. And again. Six tokens, six full sweeps through 140 GB. The arithmetic involved per token is small; the reading is what kills you. This is the memory-bandwidth wall, and it’s why your single-stream inference is slow even on a card that brags about petaflops.
Here’s the part that should really bother you: while all that reading is happening, the GPU’s compute units are mostly idle. You paid for petaflops. On dense, low-batch decode you’re using a small fraction of them. The hardware is begging you to give it more arithmetic to do per byte read.
Speculative decoding is the move that takes the hardware up on that offer.
You probably assume the catch is quality: a small model helps out, so you must be getting a slightly cheaper, slightly worse answer. It’s the opposite. The output distribution is provably identical to running the big model alone — the small model never gets a vote it can’t be overruled on. The thing you’re trading away isn’t accuracy; it’s idle silicon.
Why it matters now
The major open inference stacks ship it as a documented feature: vLLM, TensorRT-LLM, and llama.cpp all have it. Whether a given hosted API uses it is mostly not disclosed, so I won’t claim a list. For agent workloads — long, decode-heavy chains of tool calls and reasoning — it’s one of the few ways to cut wall-clock latency without dropping to a smaller model and eating the quality loss.
It also matters because it composes (mostly — real stacks have feature-by-feature gaps) with the other tricks. MoE, KV-cache reuse, prompt caching — speculative decoding stacks on top of them.
The short answer
speculative decoding = small draft model + big model verifies in parallel
Picture to keep: a junior drafting a sentence on a whiteboard while the expert reads over their shoulder — the expert checks the whole line in one glance instead of dictating word by word, crosses out from the first mistake onward, and writes the next word themselves. The expert’s standards never drop; they just stopped doing the typing.
A cheap “draft” model guesses the next K tokens. The big “target” model then runs one forward pass that scores all K guesses simultaneously, accepts the longest prefix it agrees with, and produces one bonus token of its own. If the draft is right most of the time, you get several tokens per expensive forward pass instead of one — and the output distribution is provably the same as what the big model would have produced alone.
How it works
The mechanism rests on one asymmetry that most people miss the first time they hear about this technique: in the memory-bound regime, and for small K, scoring K tokens in parallel costs the big model roughly what scoring one costs.
That’s because, in the memory-bound regime, the cost of one decode step is dominated by streaming the weights through the memory hierarchy, not by the matmuls themselves. Once a weight tile is loaded into the GPU’s caches it can be reused against a length-1 input or a length-K input at roughly the same wall-clock cost, until K gets large enough that you finally become compute-bound. So the big model can verify “did I agree with all K of these guesses?” in basically one forward pass.
Now build the technique by breaking things. Keep our sentence in mind: the 70B model is producing “The capital of France is Paris.”
Naive attempt: just use the small model. A 1B sibling gets this sentence right — it’s easy — and it’s dramatically cheaper to run, since it has ~70× fewer weights to stream per token. Ship that.
Why it breaks: you didn’t want a 1B answer, you wanted a 70B answer. On this sentence the two agree; on the hard sentence you actually care about they won’t, and from the outside you cannot tell which case you’re in. A speedup you can’t audit isn’t a speedup, it’s a downgrade.
Fix 1: let the big model check the small model’s work. The 1B model drafts K tokens ahead — say capital of France is — and the 70B model runs one forward pass over the prompt plus those four drafted tokens. That single pass hands you the target’s probability distribution at every one of those positions at once, because of the asymmetry above. Now you can ask, position by position, “would I have said that?”
Why that breaks: the draft will eventually be wrong. Suppose it drafts capital of France is Lyon. You can’t accept the whole block, and throwing the whole block away wastes the three tokens that were fine.
Fix 2: accept the longest prefix, then correct. Walk left to right and take the run of tokens the target agrees with; stop at the first disagreement. And you get one extra token essentially free at the stopping point, out of the same forward pass — if all K were accepted, you sample from the target’s distribution at position K+1, which the verify pass already computed; if you rejected at position j, you sample the replacement token from the target’s (corrected) distribution at j. Either way the big model’s pass produced somewhere between 1 and K+1 tokens instead of exactly 1. Then loop.
Why that breaks — the subtle one: “agrees with” is easy to define for greedy decoding (accept the draft token iff it’s also the target’s argmax) but not for sampling. If you sample at temperature and accept whenever the draft’s token happens to match, you have quietly reweighted the distribution toward whatever the small model likes, and your output is no longer the big model’s output. This is exactly the quality leak you were trying to avoid.
Fix 3: a rejection-sampling rule that’s provably unbiased. From the Leviathan et al. paper: accept the drafted token with probability min(1, p_target(x) / p_draft(x)), and on rejection sample from the residual (p_target − p_draft)+, normalized. The joint distribution over generated sequences is provably the target’s.
That last bit is the part that surprises engineers, and it’s the payoff for the whole chain. You’re not approximating the big model. You’re not trading quality for speed. The output distribution is identical to what the target would have produced alone. (Implementations can still differ by tiny amounts due to floating-point numerics — the guarantee is at the math level, not at the bit level.) The draft model is just a guess source; the target retains full veto power.
This is where the over-the-shoulder analogy earns its keep — and where it breaks. A real expert reading a junior’s draft is influenced by it; anchoring is a real thing. The target model isn’t. The accept/reject rule is constructed so that the draft can change how fast you get the answer and never which answer you get. That’s the difference between a helpful colleague and a provably-unbiased one.
Where the speedup comes from
Two ways to think about it, both useful:
- Compute side: you’ve converted a memory-bound workload into a slightly more compute-bound one. Each expensive forward pass over the big model now produces somewhere between 1 and K+1 tokens of output instead of exactly 1. If each drafted token is accepted with probability ~0.7 independently and K=4, the expected number of tokens per big-model step is
(1 − p^(K+1)) / (1 − p)≈ 2.8 — the accepted prefix plus the bonus token. Note how fast that saturates: pushing K from 4 to 8 at the same acceptance rate only takes you to ≈ 3.2, because a single rejection ends the run. That’s an upper bound on per-step gain — the end-to-end speedup is smaller after you pay for running the draft. - Hardware side: those previously idle compute units now have work to do — verifying the K drafted positions in parallel. You stopped wasting flops.
The two failure modes are symmetric. If the draft is too bad, acceptance rates collapse, and you’re paying for the small model’s forward passes plus the big model’s, with little to show. If the draft is too good (e.g. a 7B drafting for a 70B), the small model itself is now slow, and the cost of running it eats into the savings. Picking a draft is a Goldilocks problem.
A real-world acceptance rate of 60–80% on natural prose is typical with a well-matched draft, giving end-to-end speedups of roughly 2–3× on memory-bound inference. Reported numbers vary widely by workload — code, structured output, and chat all behave differently — and there is no canonical public benchmark to point at, so treat “2–3×” as a rough order of magnitude rather than a guarantee.
Where it stops working
Speculative decoding helps memory-bound inference. If you’re already compute-bound — large batch sizes serving many users in parallel, or very small models where the weights aren’t the bottleneck — there are no idle flops to recover, and the math stops being free. This is why hosted providers sometimes turn it on for low-batch / low-traffic regimes and off when traffic is heavy. The exact crossover point depends on the model, the hardware, and the batch size, and no clean public number exists for where modern inference servers flip the switch.
You started with speculative decoding = small draft model + big model verifies in parallel. What did the failure chain add? — + an unbiased accept/reject rule, and a hardware regime where it pays. The accept/reject rule is what turns “a cheaper approximation” into “the same answer, sooner”; the memory-bound regime is what makes the verification pass nearly free. Take away either one and the trick stops being a trick: without the rule you’ve just got a worse model, and without idle flops you’ve just got two models to pay for.
Which answers the thing we opened with. That word-by-word pace never was deliberation — it was a memory bus being read end to end, once per word. Speculative decoding doesn’t make the model think faster. It just stops making it re-read the whole library to write each one.
Check yourself
Before you go — you swap your draft model for a bigger, better one. Acceptance rate climbs from 70% to 90%. Is the system necessarily faster? What would you actually compare?
Answer
Not necessarily, and this is the Goldilocks problem from above. Higher acceptance means more accepted tokens per expensive target pass, which is pure win on the target side. But a bigger draft model costs more per drafted token, and you pay that cost K times per cycle whether the tokens get accepted or not. The comparison that matters is total wall-clock per output token: (cost of K draft steps + cost of 1 target pass) / (expected tokens accepted per cycle). If your draft went from 1B to 7B against a 70B target, you multiplied the draft-side cost roughly sevenfold to buy maybe 1.3× more tokens per cycle — likely a loss. The extreme case makes it obvious: a draft model identical to the target has 100% acceptance and zero speedup.
And one more — your provider’s docs say speculative decoding is enabled, but you see no latency improvement at all during business hours and a clear one at 3am. Nothing about your prompt changed. What’s the most likely explanation?
Answer
You’re not in the memory-bound regime during the day. Speculative decoding recovers idle compute, and at high concurrency the server is already filling those flops with other users’ requests via batching — the GPU is closer to compute-bound, so there’s nothing to recover and verifying K positions genuinely costs more than verifying one. At 3am your request is nearly alone on the card, decode is bandwidth-starved, and the free capacity is there to exploit. This is also why the technique sits awkwardly next to continuous batching: both are after the same idle silicon, and on a busy server the batching is already claiming it.
Famous related terms
- Draft model —
draft model = small LLM same tokenizer as target— a smaller sibling of the target whose job is just to propose. Often a distilled version of the target, or a much smaller model from the same family. - Medusa — adds extra “decoding heads” to the target model itself so it predicts several future tokens at once; no separate draft model.
- EAGLE — drafts at the level of the target’s own feature representations rather than over raw tokens, extrapolating the target’s hidden states forward and decoding from those. Different mechanism from Medusa, similar goal.
- Self-speculative decoding — uses earlier layers of the target as the “draft,” with later layers as the verifier. One model, two roles.
- Lookahead decoding — generates draft tokens via Jacobi iteration over the target’s own forward passes, no draft model required. Different mechanism, similar goal.
- Prompt / KV caching —
prompt cache = KV-cache + reuse across requests— orthogonal trick, also goes after the cost of re-reading work the model already did. Stacks with speculative decoding. - Continuous batching —
continuous batching = dynamic batch + per-token scheduling— the other big inference-serving win. It improves utilization and pushes the system toward compute-bound at high concurrency. When it dominates, speculative decoding’s value shrinks.
Going deeper
- Leviathan, Kalman, Matias — Fast Inference from Transformers via Speculative Decoding (arXiv 2022, ICML 2023) — the primary source, and the place to go if you want to see why the accept/reject rule provably preserves the target distribution rather than take my word for it.
- Chen, Borgeaud, Irving, Lespiau, Sifre, Jumper — Accelerating Large Language Model Decoding with Speculative Sampling (DeepMind, 2023) — concurrent independent work; read it to see the same idea derived a second way, which is decent evidence the framing isn’t an artifact of one presentation.
- NVIDIA — An Introduction to Speculative Decoding for Reducing Latency in AI Inference — the explainer: answers “what does this actually look like when I turn it on,” with diagrams and deployment framing.
- Looking back at speculative decoding — Google Research blog — the rabbit hole, for the question of how a technique gets adopted and what its authors did and didn’t anticipate.