Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why model distillation exists

A small model trained on a big model's outputs often beats the same small model trained on the original labels. That shouldn't be obvious — and the reason it works is the actually interesting part.

AI & ML intermediate Apr 29, 2026 · updated Aug 25, 2026 · 14 min read

On this page

The picture version

Six pictures for a reader who has never wondered where a small model got its abilities. The prose below fills in the seams the pictures skip.

1 · The problem

Seven billion parameters just solved a competition maths problem.

7B on your laptop a tiny fraction of the parameters answers “Let the roots be r and s. Then by Vieta’s formulas r + s = −b/a, so…” step by step, and correct ? the naive bound says a smaller network should just be a worse version of the same thing So where did the ability come from? for the cases where the training story is public, the same shape keeps appearing
The small model learned from the outputs of a much larger one. The practice is called knowledge distillation, and the thing that should surprise you is not that it saves money — it is that the student can end up better than the same small architecture trained on the original ground-truth labels.

2 · What the hard label throws away

A tick tells you the answer. A ranking tells you the subject.

the dataset’s label cat 1.0 dog 0.0 fox 0.0 truck 0.0 every wrong answer is equally wrong the teacher’s distribution cat 0.90 dog 0.07 fox 0.02 truck 0.001 dog is 70× more plausible than truck — and that ratio is real information The map of what is almost right is the actual teaching. to make those small numbers usable you raise the softmax temperature at both teacher and student, which flattens the distribution and exposes the structure among the wrong answers
A hard label says what is correct; the teacher’s distribution says what is nearly correct. That extra structure — the shape of the teacher’s uncertainty — is what a one-hot vector erases, and it is the original reason distillation beat training on the labels directly.

3 · Why the small model can keep up

It never had to discover the answer. It was shown one.

trained on the original labels noisy, finite, one bit of signal per example it has to find a good solution by itself trained on the teacher a denser, smoother signal — on far more inputs, because you can label as much as you can run inference on an easier optimisation problem on a richer dataset The bottleneck was rarely raw parameter count. it was finding the right function during training — and the teacher already did that hard part. this is the standard explanation, not a proven theorem.
Shape of the argument, not measured curves. The 7B model on your laptop did not have to work out how to solve a competition maths problem from scratch. It was shown 800,000 worked solutions by something that already could.

4 · Three things wear the same name

And only one of them needs to see inside the teacher.

1 · logit distillation match the teacher’s full next-token distribution needs the logits the 2015 recipe. DistilBERT. you must own the teacher 2 · sequence distillation let the teacher write, then fine-tune on what it wrote needs only its samples the R1-Distill checkpoints: ~800k traces, plain fine-tuning, no RL works through a closed API 3 · synthetic-data pretraining a good model generates filtered training data for somebody else’s base is this still distillation? the line gets blurry here — terminology drift, not a clean technical boundary Your laptop model came from the middle column, not the left one. which is why the compression line has to loosen: what actually transferred was the teacher’s behaviour, never its probabilities
Every sample collapses the teacher’s distribution to one draw, so column 2 gives up exactly the “dark knowledge” that made column 1 work — and needs far more samples to recover comparable information. It is also the only column available to anyone who can query a model but not inspect it.

5 · The seam

Human errors are noise. A teacher’s errors are signal.

human-labelled data labeller A → “cat” labeller B → “fox” labeller C → “cat” the model sees disagreement — errors look like noise one teacher’s outputs teacher → “fox” teacher → “fox” teacher → “fox” no disagreement anywhere — the mistake is now a clean label The student inherits the teacher’s blind spots, and cannot see them. often more so, because it has less capacity to hold a contrarian view it never saw in training which is why “the student scores within a couple of points on our benchmark” is not the same as “we can retire the teacher”
This is the failure mode you cannot audit easily: the student’s mistakes are correlated with the teacher’s rather than scattered. Distillation transfers a teacher’s answers, not its judgement — and the teacher was also the labelling machine you would need to handle anything new.

6 · Keep this card

The whole thing on one index card.

distillation ≈ train a small student on a big model’s behaviour + the full distribution is the richest form of that    — and only available if you own the teacher ∴ 7B doesn’t secretly suffice — someone paid for 671B first
Picture to keep: a marked exam paper where the teacher didn’t just circle the right answer but ranked all the wrong ones — “this one was nearly right, that one was nowhere near.” Where the analogy breaks: a human teacher marks a fixed exam, whereas you can run a teacher model over as much unlabelled input as you can afford. The volume matters as much as the richness.

Why it exists

At some point you’ve noticed the cheap option is suspiciously good. The “mini” tier of an API costs a fraction of the flagship and answers most of your questions just as well; a 7B model running on your own laptop walks through a competition maths problem you’d have bet only a datacentre could do. Something doesn’t add up — the small model has a tiny fraction of the parameters, so where did the ability come from?

For the cases where the training story is public, the same shape keeps appearing: at some stage, the small model learned from the outputs of a much larger model. DeepSeek-R1-Distill-Qwen-7B is the cleanest public example and it’s the running example for this post. The practice has a name: knowledge distillation.

The thing that should surprise you is that it works at all. The naive intuition is: a smaller network has less capacity, so the best you can hope for is “trains on the same data, ends up a bit worse.” Distillation ignores that bound. It says: take a giant teacher model, run it on a pile of inputs, record the whole probability distribution it produces (or, more recently, the actual generated text), and train a small student to imitate that. The student often ends up better than the same small architecture trained on the original ground-truth labels. That’s the trick — and it’s been re-discovered, generalized, and re-tooled three times across a decade.

The earliest version usually credited is Bucilă, Caruana and Niculescu-Mizil’s 2006 Model Compression paper, which trained small neural nets to mimic big ensembles by labelling a large unlabelled set with the ensemble’s predictions. The version most people mean today is Hinton, Vinyals and Dean’s 2015 Distilling the Knowledge in a Neural Network, which gave us the modern recipe: soft targets from a softened softmax. The version your inference bill cares about in 2026 is small dense models distilled from frontier reasoning models on synthetic data.

Why it matters now

For an engineer shipping with LLMs, distillation is a large part of why “small enough to be cheap” and “good enough to be useful” overlap at all:

The pragmatic version of the question, then, is “when can I get away with the small one?” — and to answer that you have to know what distillation actually transfers and what it doesn’t.

The short answer

distillation = train a small student on the teacher's full output distribution + (optionally) the original labels

Picture to keep: a marked exam paper where the teacher didn’t just circle the right answer but ranked all the wrong ones — “this one was nearly right, that one was nowhere near.” A hard label tells the student what’s correct; the teacher’s distribution tells it what’s almost correct, and the almost-correct map is the actual teaching. Where the analogy breaks: a human teacher marks a fixed exam, whereas you can run the teacher model over as much unlabelled input as you can afford — the volume matters as much as the richness.

Instead of teaching the student “the answer is class 7” with a one-hot label, you teach it “the teacher thought class 7 was 62% likely, class 3 was 18%, class 9 was 11%, …”. That extra structure — the shape of the teacher’s uncertainty — turns out to carry a lot of information that a hard label throws away.

For modern LLMs the same idea shows up in two flavours: matching the teacher’s per-token probability distribution (true distillation), or just training on text the teacher generated (often called distillation loosely; technically closer to “synthetic-data fine-tuning”). The two get conflated in casual usage, and the distinction matters in places.

How it works

The original trick: soft targets

Start with the simplest case — a classifier — and then we’ll come back to the 7B model on your laptop. The standard training signal for an input is a one-hot vector: the correct class is 1.0, everything else is 0.0. Hinton’s 2015 paper observed something that, in retrospect, is obvious. The teacher’s softmax doesn’t just say “cat.” It says something like cat 0.90, dog 0.07, fox 0.02, truck 0.001. The ratio of “dog” to “truck” — both wrong — encodes real similarity information that the dataset’s hard label has erased. The phrase dark knowledge gets attached to this idea — it’s associated with Hinton from later talks rather than the 2015 paper itself, but it’s the label most people use.

To make those small numbers usable, you raise the softmax temperature T at both teacher and student during training:

p_i = exp(z_i / T) / Σ_j exp(z_j / T)

A higher T flattens the distribution and exposes the structure in the “wrong” classes. Hinton et al. report using temperatures from about 1 up to 20 in their experiments. The student is trained with a weighted sum of two losses: cross-entropy with the soft targets at high T, plus ordinary cross-entropy with the true labels at T = 1. After training, the student runs at T = 1 like any other model. (Hinton et al., 2015)

Why a small model can match a big one (sometimes)

The capacity argument — “smaller networks must be worse” — is misleading. The real bottleneck for a small model is rarely raw parameter count; more often it’s finding the right function during training. A big model trained on noisy, finite labelled data has done the hard work of locating a good decision surface. Distilling it gives the student an effectively denser, smoother training signal: more nuanced labels, often on far more inputs (you can label as much unlabelled data as you can run inference on). The student is solving an easier optimization problem on a richer dataset. That’s the answer to the puzzle you opened with: the 7B model on your laptop didn’t have to discover how to work through a competition maths problem. It was shown 800k worked solutions by something that already could.

This is exactly what Bucilă, Caruana and Niculescu-Mizil showed in 2006, before the deep-learning era: small neural nets trained to mimic the output of a large ensemble matched the ensemble’s accuracy while being, in their conclusion’s wording, roughly 1000× smaller and faster on average. The “more inputs, ensemble-labelled” part of their recipe is the part the modern LLM era leans on hardest. (Bucilă et al., 2006)

Distillation in the LLM era: three things that get called the same name

Here’s where the term gets fuzzy. In current practice, “distillation” covers at least three setups:

  1. Logit / soft-target distillation. The classical Hinton recipe, adapted to LLMs: at every position, train the student to match the teacher’s full next-token distribution (or its top-k). Requires access to the teacher’s logits. DistilBERT (Sanh et al., 2019) is a well-known example — about 40% smaller than BERT-base, retaining roughly 97% of its GLUE performance, around 60% faster at inference. (Sanh et al., 2019)
  2. Sequence / behavior distillation. Have the teacher generate text on a set of prompts, then fine-tune the student via ordinary supervised fine-tuning on those (prompt, teacher-output) pairs. This is what the R1-Distill checkpoints — including the 7B one on your laptop — used: ~800k reasoning traces from R1, plain SFT on the smaller bases, no RL. You only need the teacher’s samples, not its logits — which is why this works through a closed-API teacher.
  3. Self-distillation and synthetic-data pretraining. The teacher and student can be the same model architecture at different training stages, or the teacher can simply be a high-quality model used to generate filtered pretraining-style data for somebody else’s base. The line between “distillation” and “training on synthetic data” gets blurry here. I’d flag it as terminology drift in the field rather than a clean technical boundary.

The first kind is what the 2015 paper meant. The second kind is what the public DeepSeek-R1 example does. The third kind appears to be common practice in the current small-model ecosystem, but quantifying “how common” runs straight into the wall of closed-lab opacity.

Where the seams are

A few things worth knowing if you ever lean on a distilled model in production:

You started with distillation = train a small student on the teacher's full output distribution. Notice what the DeepSeek example quietly broke — it never touched R1’s probabilities at all, only its generated text. So the honest line is distillation ≈ train a small student on a big model's *behaviour*, and “the full distribution” is just the richest form of that behaviour you can get when you own the teacher. That’s why your laptop model can reason: not because 7B parameters secretly suffice, but because someone paid for 671B parameters to produce 800k worked examples first.

Check yourself

Before you go — you distill a frontier model into a 7B student, and the student scores within a couple of points of the teacher on your benchmark. Your PM wants to drop the expensive teacher entirely. What’s the strongest technical objection you can raise?

Answer

That the benchmark measures the region the student was taught, and distillation transfers a teacher’s answers rather than its judgement. Two specific worries: first, the student inherits the teacher’s confident errors as clean training signal, with no disagreement in the data to soften them — so its mistakes will be correlated with the teacher’s and harder to catch than noisy human-label errors. Second, you now have no way to generate new training data or to handle distribution shift; the teacher was the labelling machine. Matching on today’s benchmark is not the same as being able to keep up.

And one more: a startup wants to distill from a closed API model they can only query, not inspect. Which of the three flavours in this post is available to them, and what do they lose?

Answer

Only sequence/behaviour distillation — fine-tuning on (prompt, teacher-output) pairs — because logit distillation needs the teacher’s per-token probabilities and an API returns sampled text. What they lose is exactly the “dark knowledge”: the ranking of the wrong answers. Every sample collapses the teacher’s distribution to one draw, so they need far more samples to recover comparable information. (Separately, and more importantly in practice: check the provider’s terms — several restrict using outputs to train competing models.)

Going deeper

Well-established: the mechanical story — soft targets carry more information than hard labels, and sequence distillation works through teacher samples alone. Not public: the relative contribution of distillation vs. base-data quality vs. architecture choices in any specific small-model success story. The closed labs don’t publish the breakdown, and reverse-engineering it from leaderboards is unreliable, so this post leans on the open examples where the recipe is documented.