Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why do small models exist?

If bigger models always benchmark better, why does anyone ship a 3B model? The answer is mostly about latency, cost, and the place the model has to live.

AI & ML intro Apr 29, 2026 · updated Aug 25, 2026 · 9 min read

On this page

The picture version

Six pictures for a reader who assumes bigger is simply better. The prose below fills in the seams the pictures skip.

1 · The thing you’ve noticed

Same laptop, same app, two completely different speeds.

the grey ghost-text con st total = items.reduce( this gap — two keystrokes apart late is the same as missing the chat panel “why is this function slow?” several seconds — and that’s fine you asked, so you’ll wait These are almost never the same model.
One feature has a couple of hundred milliseconds before it becomes an irritation; the other can take its time. Two jobs on one machine, with budgets that differ by an order of magnitude.

2 · The missing denominator

Leaderboards rank answers. Products pay per millisecond.

the leaderboard the big model 78 the small model 73 one column. quality, and nothing else. the columns products care about per millisecond per dollar per watt per device it fits on small wins small wins small wins small wins five points behind, and five times faster A suggestion that arrives too late isn’t a worse suggestion. It’s no suggestion.
The ranking only flips when you put the denominators back. For a large share of real work, the question isn’t which model answers best — it’s which model answers best inside the budget the job actually has.

3 · Why smaller is faster than you’d guess

Speed is set by how many bytes have to be read per word.

140 GB read for every single word a 70-billion-parameter model ~1.5 GB a 3-billion-parameter model, stored at four bits per number While the model writes, the chip’s arithmetic units mostly wait for bytes. so cutting the model’s size cuts the wait almost proportionally — that’s where the snappiness comes from
Generating text is limited by moving the model’s numbers, not by multiplying them. Fewer and smaller numbers means less to move, which is why the speed gap between a big and a small model is closer to the gap in bytes than to the gap in benchmark scores.

4 · The slot

Some places have a hard ceiling. Nothing bigger goes in.

a phone a laptop a browser tab fits fits fits hard ceilings on memory, power and heat — not negotiable the frontier model: a rack of accelerators, working together can’t go here no network · no per-word bill · nothing leaves the device
The other half of the answer isn’t speed at all — it’s that the biggest models physically cannot run where a lot of the work happens. If you want the capability there, the model has to fit there, and that changes what the product can promise.

5 · Where it stops working

Fast, cheap, always on-call — and thin on the hard tail.

the requests the pattern-shaped majority the hard tail the small model instant · cheap · local the big model slower · pricier · remote obscure facts long chains of reasoning an unfamiliar codebase problems nobody has shaped before ← what the small one is thin on Small by default, escalate on demand — a product decision, not a ranking.
Think of a junior who is fast, cheap and always available — excellent on the routine majority, weak where the work needs deep recall or genuine novelty. The analogy breaks in one place: a junior learns your codebase over time, and the model’s numbers are frozen.

6 · Keep this card

The whole idea on one index card.

small model = fewer numbers + the same modern training recipe + a slot the big model can’t fit into — the slot is what picks the model, not the leaderboard
Picture to keep: the autocomplete model lives inside the gap between two of your keystrokes; the big model lives in a datacentre and has to mail its answer back.

Why it exists

You type three characters in your editor and grey ghost-text finishes the line before your next keystroke lands. Then you ask the chat panel in the same editor a question, and you watch it think for several seconds. Same laptop, same vendor, wildly different feel — because those two features are almost never the same model. The autocomplete has a couple hundred milliseconds, give or take, before it’s more annoying than useful (that’s a product rule of thumb, not a published number); the chat answer can afford to take its time. Keep that inline autocomplete in mind — it’s the running example for the whole post.

Open any leaderboard and the pattern looks like it should rule out that first model existing at all: hold the architecture and training recipe roughly constant, scale parameters up, scores go up. Scaling has been the dominant story of the last several years of LLM progress. So why do labs keep shipping tiny siblings — 1B, 3B, 7B models — alongside the flagship?

Because leaderboards rank on answer quality and nothing else. The autocomplete is graded on quality per millisecond, per dollar, per watt, per device. Put those denominators back in and the ranking flips for a huge fraction of real workloads: a model that scores five points lower but answers in a fifth of the time wins the autocomplete slot outright, because a slow suggestion isn’t a worse suggestion — it’s no suggestion at all.

There’s a second reason that’s easy to miss: the biggest models can’t physically run where the work is. The largest frontier models are served from racks of accelerators, each carrying tens of gigabytes of fast memory, working together on one model. A laptop, a phone, a browser tab, an edge gateway, a car — these places have hard ceilings on RAM, power, and thermals. If you want intelligence there, the model has to fit there.

Why it matters now

Three pressures put the small-model question in front of anyone shipping a product:

  1. Inference is where the bill lives. Training is a one-time spike; serving is a cost you pay forever, once per request. So a model that’s much cheaper and much faster per call, and loses only a few points on your task, is often the one that survives the budget review.
  2. Agents fan out. An agentic loop can call the model tens or hundreds of times for a single task. Multiply per-call latency and per-call cost by that, and you understand why teams reach for the smallest model that still does the step.
  3. On-device is no longer a toy. Phones ship neural accelerators, laptops ship unified memory, and browsers expose WebGPU. A 3B model’s weights stored at 4 bits each work out around 1.5 GB — call it a couple of gigabytes once you add the runtime and the KV cache — which makes it a different product category from a cloud API: no network, no per-token cost, no data leaving the device. The editor autocomplete and the phone’s on-device text features live here.

The short answer

small model = fewer parameters + the same modern recipe + a deployment slot the big model can't fit into

Picture to keep: the autocomplete model lives inside the latency budget between two of your keystrokes; the big model lives in a datacenter and has to mail its answer back.

A small model isn’t a worse big model — it’s a model deliberately built to fit a budget (memory, latency, power, dollars) where a big model can’t run at all, or can’t run economically. “The same modern recipe” is the important half of that line: today’s small models get the same training treatment as the flagships — the same kind of curated data, the same instruction and preference tuning — rather than being an older, cruder thing that happens to be small. You trade some peak capability for the ability to actually exist in that slot.

How it works

Start with the naive move and watch it break.

Naive attempt: just call the frontier model for the autocomplete too. It’s the smartest model you have. But a round trip to a datacenter plus a forward pass through a frontier-scale model doesn’t come back inside a keystroke gap, and you’d be paying frontier prices for a suggestion the user discards most of the time. On a laptop in airplane mode it can’t run at all: the weights don’t fit in the RAM you have.

Fix: shrink the model. Fewer parameters means less to read and less to compute. The size of the win is bigger than the FLOP count suggests, because decoding is typically memory-bandwidth-bound: the accelerator’s arithmetic units finish early and wait on bytes arriving from memory. Halve the parameters and you roughly halve the weight bytes read per token — which is why the gap between a 7B and a 70B on the same hardware tracks the tenfold difference in bytes, not the smaller difference in raw arithmetic. (The memory bandwidth post has the gory details, including the KV cache’s share of that traffic.)

But a small model trained the old way isn’t good enough to ship. Parameter count isn’t the only thing that moved: the recipe upgrades that made frontier models smarter — better data mixes, instruction tuning, preference optimization — apply at every size. A model with 3 billion parameters today is a meaningfully different product from one with 3 billion parameters two generations ago. How much better is hard to state as a single number: the answer depends on which benchmark and which model family you compare, and cross-family comparisons mix in differences in training data nobody publishes.

And a small model trained on raw web text alone leaves capability on the table. A cheaper route is to have a frontier model generate a large pile of high-quality outputs and train the small one to imitate them — distillation. The student can’t match the teacher, but on the slice of behaviour it was distilled on, it gets close. That’s the standard account of why a lab’s “small” model feels more competent than its size suggests; how much of any specific model’s quality comes from distillation versus data curation is not something labs disclose.

But the small model still falls over on the hard tail. Obscure facts, multi-step reasoning under pressure, novel problem decomposition, code on an unusual stack. The mental image that holds up is a junior who is fast, cheap, and always on-call — great for the pattern-shaped majority of the work, bad for the part that needs taste or deep memory. The analogy breaks in one specific place: a junior gets better as they see more of your codebase, and the small model’s weights are frozen at inference time. It doesn’t learn from yesterday’s session. A product can hand it notes — retrieved files, saved preferences, a later round of fine-tuning — but that’s the surrounding system remembering, not the model.

Fix: don’t ask it to do the hard tail. Route the easy majority to the small model and escalate the rest to a big one — the cascade pattern. That’s the answer to the thing you noticed in your editor: the ghost-text and the chat panel are two different models on the same laptop because they were selected by two different budgets, and the routing between them is a product decision, not a capability one.

You started with small model = fewer parameters + the same modern recipe. What did this post add? — + a deployment slot, and that’s the load-bearing part: the slot (a keystroke gap, a couple of gigabytes of phone RAM, a per-call cost multiplied by a hundred agent steps) is what selects the model, not the leaderboard. Try it on a case this post didn’t cover — a phone dictating a voice memo into cleaned-up text with no signal — and the slot tells you the answer before any benchmark does.

Going deeper