Why do small models exist?
If bigger models always benchmark better, why does anyone ship a 3B model? The answer is mostly about latency, cost, and the place the model has to live.
On this page
The picture version
Six pictures for a reader who assumes bigger is simply better. The prose below fills in the seams the pictures skip.
1 · The thing you’ve noticed
Same laptop, same app, two completely different speeds.
2 · The missing denominator
Leaderboards rank answers. Products pay per millisecond.
3 · Why smaller is faster than you’d guess
Speed is set by how many bytes have to be read per word.
4 · The slot
Some places have a hard ceiling. Nothing bigger goes in.
5 · Where it stops working
Fast, cheap, always on-call — and thin on the hard tail.
6 · Keep this card
The whole idea on one index card.
Why it exists
You type three characters in your editor and grey ghost-text finishes the line before your next keystroke lands. Then you ask the chat panel in the same editor a question, and you watch it think for several seconds. Same laptop, same vendor, wildly different feel — because those two features are almost never the same model. The autocomplete has a couple hundred milliseconds, give or take, before it’s more annoying than useful (that’s a product rule of thumb, not a published number); the chat answer can afford to take its time. Keep that inline autocomplete in mind — it’s the running example for the whole post.
Open any leaderboard and the pattern looks like it should rule out that first model existing at all: hold the architecture and training recipe roughly constant, scale parameters up, scores go up. Scaling has been the dominant story of the last several years of LLM progress. So why do labs keep shipping tiny siblings — 1B, 3B, 7B models — alongside the flagship?
Because leaderboards rank on answer quality and nothing else. The autocomplete is graded on quality per millisecond, per dollar, per watt, per device. Put those denominators back in and the ranking flips for a huge fraction of real workloads: a model that scores five points lower but answers in a fifth of the time wins the autocomplete slot outright, because a slow suggestion isn’t a worse suggestion — it’s no suggestion at all.
There’s a second reason that’s easy to miss: the biggest models can’t physically run where the work is. The largest frontier models are served from racks of accelerators, each carrying tens of gigabytes of fast memory, working together on one model. A laptop, a phone, a browser tab, an edge gateway, a car — these places have hard ceilings on RAM, power, and thermals. If you want intelligence there, the model has to fit there.
Why it matters now
Three pressures put the small-model question in front of anyone shipping a product:
- Inference is where the bill lives. Training is a one-time spike; serving is a cost you pay forever, once per request. So a model that’s much cheaper and much faster per call, and loses only a few points on your task, is often the one that survives the budget review.
- Agents fan out. An agentic loop can call the model tens or hundreds of times for a single task. Multiply per-call latency and per-call cost by that, and you understand why teams reach for the smallest model that still does the step.
- On-device is no longer a toy. Phones ship neural accelerators, laptops ship unified memory, and browsers expose WebGPU. A 3B model’s weights stored at 4 bits each work out around 1.5 GB — call it a couple of gigabytes once you add the runtime and the KV cache — which makes it a different product category from a cloud API: no network, no per-token cost, no data leaving the device. The editor autocomplete and the phone’s on-device text features live here.
The short answer
small model = fewer parameters + the same modern recipe + a deployment slot the big model can't fit into
Picture to keep: the autocomplete model lives inside the latency budget between two of your keystrokes; the big model lives in a datacenter and has to mail its answer back.
A small model isn’t a worse big model — it’s a model deliberately built to fit a budget (memory, latency, power, dollars) where a big model can’t run at all, or can’t run economically. “The same modern recipe” is the important half of that line: today’s small models get the same training treatment as the flagships — the same kind of curated data, the same instruction and preference tuning — rather than being an older, cruder thing that happens to be small. You trade some peak capability for the ability to actually exist in that slot.
How it works
Start with the naive move and watch it break.
Naive attempt: just call the frontier model for the autocomplete too. It’s the smartest model you have. But a round trip to a datacenter plus a forward pass through a frontier-scale model doesn’t come back inside a keystroke gap, and you’d be paying frontier prices for a suggestion the user discards most of the time. On a laptop in airplane mode it can’t run at all: the weights don’t fit in the RAM you have.
Fix: shrink the model. Fewer parameters means less to read and less to compute. The size of the win is bigger than the FLOP count suggests, because decoding is typically memory-bandwidth-bound: the accelerator’s arithmetic units finish early and wait on bytes arriving from memory. Halve the parameters and you roughly halve the weight bytes read per token — which is why the gap between a 7B and a 70B on the same hardware tracks the tenfold difference in bytes, not the smaller difference in raw arithmetic. (The memory bandwidth post has the gory details, including the KV cache’s share of that traffic.)
But a small model trained the old way isn’t good enough to ship. Parameter count isn’t the only thing that moved: the recipe upgrades that made frontier models smarter — better data mixes, instruction tuning, preference optimization — apply at every size. A model with 3 billion parameters today is a meaningfully different product from one with 3 billion parameters two generations ago. How much better is hard to state as a single number: the answer depends on which benchmark and which model family you compare, and cross-family comparisons mix in differences in training data nobody publishes.
And a small model trained on raw web text alone leaves capability on the table. A cheaper route is to have a frontier model generate a large pile of high-quality outputs and train the small one to imitate them — distillation. The student can’t match the teacher, but on the slice of behaviour it was distilled on, it gets close. That’s the standard account of why a lab’s “small” model feels more competent than its size suggests; how much of any specific model’s quality comes from distillation versus data curation is not something labs disclose.
But the small model still falls over on the hard tail. Obscure facts, multi-step reasoning under pressure, novel problem decomposition, code on an unusual stack. The mental image that holds up is a junior who is fast, cheap, and always on-call — great for the pattern-shaped majority of the work, bad for the part that needs taste or deep memory. The analogy breaks in one specific place: a junior gets better as they see more of your codebase, and the small model’s weights are frozen at inference time. It doesn’t learn from yesterday’s session. A product can hand it notes — retrieved files, saved preferences, a later round of fine-tuning — but that’s the surrounding system remembering, not the model.
Fix: don’t ask it to do the hard tail. Route the easy majority to the small model and escalate the rest to a big one — the cascade pattern. That’s the answer to the thing you noticed in your editor: the ghost-text and the chat panel are two different models on the same laptop because they were selected by two different budgets, and the routing between them is a product decision, not a capability one.
You started with small model = fewer parameters + the same modern recipe. What did this post add? — + a deployment slot, and that’s the load-bearing part: the slot (a keystroke gap, a couple of gigabytes of phone RAM, a per-call cost multiplied by a hundred agent steps) is what selects the model, not the leaderboard. Try it on a case this post didn’t cover — a phone dictating a voice memo into cleaned-up text with no signal — and the slot tells you the answer before any benchmark does.
Famous related terms
- Distillation —
distillation = big teacher model + small student model + imitation training— how labs squeeze frontier behavior into a phone-sized body. - Quantization —
quantization ≈ storing weights in fewer bits— turns a 7B model from ~14GB (fp16) into ~4GB (int4) so it fits in laptop RAM, with usually-small quality loss. - MoE — total parameters are huge but only a small fraction activate per token, so the effective serving cost looks small even when the model isn’t.
- Cascade / router —
cascade = small model first + escalate to big model on hard inputs— the cheap-by-default pattern that small models enable. - SLM —
SLM ≈ small model + a narrow target job— the marketing term you’ll see on vendor pages; a well-tuned one can beat a generalist big model on that specific job.
Going deeper
- Training Compute-Optimal Large Language Models (Hoffmann et al., 2022 — the “Chinchilla” paper) — read it for the question “for a fixed compute budget, how big should the model actually be?”, whose answer is why a well-trained smaller model can beat a badly-trained larger one.
- Distilling the Knowledge in a Neural Network (Hinton, Vinyals & Dean, 2015) — the primary source for “how does a small model learn from a big one?”; the technique predates LLMs entirely but is central to how today’s small models are built.
- Rabbit hole: the model card published alongside any small model you’re considering — the fastest way to see what a given small variant was actually optimized for, and which of its numbers are quality numbers versus budget numbers.