Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

What does 'X parameters' mean in an LLM?

Llama 3.1 70B, DeepSeek-V3 671B, Phi-4 14B — what is that number actually counting, and why is it the headline figure on every model release?

AI & ML intro May 4, 2026 · updated Aug 25, 2026 · 11 min read

On this page

The picture version

Six pictures for a reader who has only ever seen the number on a download page. The prose below fills in the seams the pictures skip.

1 · The problem

Three buttons, one number, no explanation.

8B 70B 405B download · ~16 GB download · ~140 GB download · much more ? your machine The “B” is for billion. Billion what, exactly? it sets the price per word, the hardware you need, and how models get ranked
The same number is the headline on every model release, the reason bigger versions cost more, and the first thing an engineer checks. It is almost never explained on the page it appears on.

2 · What it counts

Cells in a very large stack of number grids.

0.0417 one parameter a single number that training set, not a person “70B” = about 70 billion cells like that one. 80 grids deep for a 70B model · each grid thousands of columns wide Running the model = dragging every generated word past those cells.
A parameter is one learned number: nobody wrote it, and it ended up where it is because guessing the next chunk of text got better that way. The headline number is simply how many there are.

3 · The first thing that breaks

File size isn’t the count. It’s the count times the format.

140 GB 2 bytes per number 70,000,000,000 learned numbers 35 GB half a byte per number, rounded 70,000,000,000 learned numbers = Same model. Same count. One quarter the file. shrinking the storage format rounds the numbers — so it isn’t free — but it doesn’t remove a single one of them
Ask “how big is this model?” with a file size and you measure the storage format instead. The count of learned numbers is the thing underneath that doesn’t move — which is why that is what gets printed on the box.

4 · What they cost

Two bills, both straight lines in the count.

memory to hold it count × bytes each 70B × 2 bytes ≈ 140 GB this is the number that decides whether it fits on your hardware at all work per generated word ≈ 2 × count 70B → ~140 billion operations, per word every cell takes part in one multiply and one add — hence roughly two Double the count and you roughly double both bills. most of the cells sit in the per-word transformation layers, not in attention — that’s where the budget goes
This is why parameter count became the headline: it is the one figure that turns straight into both the hardware you need and the price of a word. Both relationships assume every cell is used for every word.

5 · The second thing that breaks

Some models only switch on a slice of themselves.

one word a router picks two the rest stay switched off for this word 671B on the shelf ~37B per word still needs the memory for all of it — but costs like a much smaller model to run so the shelf number is an accounting choice, not a measurement
The neat “two operations per parameter” rule quietly assumed every parameter fires. Once a model routes each word to a couple of its many parts, total size and per-word cost stop being the same story — and the headline stops being comparable across models.

6 · Keep this card

The whole thing on one index card.

parameter = one learned number the count is what the file format can’t change + cost tracks it — while every cell fires + the number on the box is an accounting choice — ask: dense or routed? total or active?
Picture to keep: a stack of 80 spreadsheets, each thousands of columns wide, and every generated word dragged past every cell — except in the models where most of the stack is skipped.

Why it exists

You decide to run a model on your own machine. The download page offers you Llama 3.1 8B, 70B, and 405B, and nothing else on the page tells you which one your laptop can survive. Same for the API pricing table, where the bigger number usually costs more per token. Same for every leaderboard — DeepSeek-V3 671B, Phi-4 14B, GPT-3 175B. The “B” — for billion — is doing real work: it’s how people compare models at a glance, why bigger SKUs cost more per token, how engineers decide whether the thing fits on their GPU. But the number is rarely explained. What is it counting? (Llama 3.1 70B is the example I’ll keep coming back to.)

It’s the parameter count: the total number of values inside the model that were learned during training, not written down by a human. When a release says “70B,” it means the model file on disk holds roughly 70 billion such numbers, each one shaped by gradients from billions or trillions of training tokens until next-token prediction got better.

The reason this number — instead of, say, lines of code or layer count — became the headline is that the two costs you care about scale directly with it: how much memory you need to load the weights, and how many floating-point operations each generated token takes. Capability tracks it too, but far more loosely.

The scaling-law work from around 2020 is what made the number feel authoritative (Kaplan et al., 2020): for LLMs built on the transformer architecture and trained on next-token prediction, the loss falls predictably as you scale, and architecture details — depth vs. width, head count, exact vocabulary — are second-order corrections inside a wide range. It’s worth naming the simplification, though: parameter count was never the only dial. Kaplan et al. and the later Chinchilla work (Hoffmann et al., 2022) disagreed about how to split a fixed compute budget between parameters and training tokens, and the second answer — train smaller models on far more data — is closer to what labs actually do now. So “N parameters” is a headline the field partly moved past; my read of why it stayed on the box is that it’s the one number that maps cleanly onto your GPU.

Why it matters now

Three places the number lands as a real engineering constraint, not a marketing figure:

The short answer

parameter = one learned number inside the model

Picture to keep: a stack of 80 spreadsheets of numbers, each about 8192 columns wide, and every generated token has to be dragged past every cell. The parameter count is how many cells there are. Except that the cells aren’t read in isolation — they’re multiplied in blocks, and in a mixture-of-experts model most of the stack is skipped for any given token, which is exactly where the headline number stops meaning what you think.

A parameter count is the total number of values (weights and biases) inside the neural network that were set by training rather than by a human writing code. “Llama 3.1 70B” means there are about 70 billion such numbers in the model file. Most of them live in matrices used by the transformer’s feed-forward and attention layers; multiplying input vectors through those matrices is most of what running the model means.

How it works

The interesting part is watching the simplest version of “how big is this model?” break, three times.

Naive attempt: just look at the file size. The 70B download is ~140 GB; the 8B one is ~16 GB. Done? Why it breaks: re-download the same 70B model quantized to 4 bits and the file is ~35 GB. The architecture didn’t change and neither did the number of learned values — only the precision each one is stored at (they’re rounded, so it isn’t free). File size mostly measures storage format. The invariant underneath is the count of learned numbers, and that’s the parameter count.

Where the parameters live

A transformer LLM is mostly stacks of matrix multiplications. For a model with hidden dimension d and L layers, each layer contains roughly:

Adding up: a “vanilla” transformer layer is roughly 12d² parameters. With L layers, the body of the model is ~12 · L · d². The embedding and output tables add ~2 · V · d where V is the vocab size; for big vocabularies (100k+) this is a couple of percent of the total.

For Llama 3.1 70B (d = 8192, L = 80, V ≈ 128k), the back-of-envelope is 12 × 80 × 8192² ≈ 64B from the body plus a couple of billion for embeddings — not exactly 70B, but close enough that the math is recognizably doing the right thing. The remaining gap comes from architecture specifics (SwiGLU’s three matrices, the precise FFN width, untied output embeddings). The full architecture table is in the Llama 3 paper; the formula above is the part worth carrying in your head.

The takeaway: the FFN matrices hold most of the parameters — typically around two-thirds of the body. This is why the FFN is what mixture-of-experts replaces. That’s where the budget is.

What the parameters cost

So now you can count them. Second naive attempt: assume the count tells you what the model costs to run. For a dense model this mostly holds, and it’s worth seeing exactly why. Two physical costs scale linearly with parameter count:

The 2N-per-token relationship is why inference cost is so legible. Doubling the parameter count roughly doubles the per-token cost. It’s also why mixture-of-experts is interesting: an MoE model only fires a subset of its parameters per token, so total params and active params come apart.

What gets counted

Why that breaks: the assumption hiding inside “N parameters ⇒ 2N FLOPs per token” is that every parameter is used for every token. Once that stops being true, the headline number stops being a cost estimate — and a few other accounting choices quietly move it too.

So when you read “X B parameters,” ask: is it dense or MoE? (If MoE: total or active?) Total or non-embedding? For most dense model cards the headline is total parameters unless stated otherwise, and the math above lines up.

You started with parameter = one learned number inside the model. What did the three breakages add? — + N is the storage-format-independent invariant, + cost ≈ 2N FLOPs/token only when every parameter fires, and + the number on the box is an accounting convention, not a measurement. That’s why “DeepSeek-V3 671B” and “Llama 3.1 70B” are only comparable on one axis: both tell you shelf size, and only the dense one also tells you roughly what a token costs.

Going deeper