What does 'X parameters' mean in an LLM?
Llama 3.1 70B, DeepSeek-V3 671B, Phi-4 14B — what is that number actually counting, and why is it the headline figure on every model release?
On this page
The picture version
Six pictures for a reader who has only ever seen the number on a download page. The prose below fills in the seams the pictures skip.
1 · The problem
Three buttons, one number, no explanation.
2 · What it counts
Cells in a very large stack of number grids.
3 · The first thing that breaks
File size isn’t the count. It’s the count times the format.
4 · What they cost
Two bills, both straight lines in the count.
5 · The second thing that breaks
Some models only switch on a slice of themselves.
6 · Keep this card
The whole thing on one index card.
Why it exists
You decide to run a model on your own machine. The download page offers you Llama 3.1 8B, 70B, and 405B, and nothing else on the page tells you which one your laptop can survive. Same for the API pricing table, where the bigger number usually costs more per token. Same for every leaderboard — DeepSeek-V3 671B, Phi-4 14B, GPT-3 175B. The “B” — for billion — is doing real work: it’s how people compare models at a glance, why bigger SKUs cost more per token, how engineers decide whether the thing fits on their GPU. But the number is rarely explained. What is it counting? (Llama 3.1 70B is the example I’ll keep coming back to.)
It’s the parameter count: the total number of values inside the model that were learned during training, not written down by a human. When a release says “70B,” it means the model file on disk holds roughly 70 billion such numbers, each one shaped by gradients from billions or trillions of training tokens until next-token prediction got better.
The reason this number — instead of, say, lines of code or layer count — became the headline is that the two costs you care about scale directly with it: how much memory you need to load the weights, and how many floating-point operations each generated token takes. Capability tracks it too, but far more loosely.
The scaling-law work from around 2020 is what made the number feel authoritative (Kaplan et al., 2020): for LLMs built on the transformer architecture and trained on next-token prediction, the loss falls predictably as you scale, and architecture details — depth vs. width, head count, exact vocabulary — are second-order corrections inside a wide range. It’s worth naming the simplification, though: parameter count was never the only dial. Kaplan et al. and the later Chinchilla work (Hoffmann et al., 2022) disagreed about how to split a fixed compute budget between parameters and training tokens, and the second answer — train smaller models on far more data — is closer to what labs actually do now. So “N parameters” is a headline the field partly moved past; my read of why it stayed on the box is that it’s the one number that maps cleanly onto your GPU.
Why it matters now
Three places the number lands as a real engineering constraint, not a marketing figure:
- Memory. Each parameter has to be stored as a number, and the format is a choice. In bf16, a common training format, that’s 2 bytes, so a 70B model needs ~140 GB just to hold the weights — too big for one 80 GB H100, so you’re into multi-GPU sharding, with its own overheads. Quantization to int4 brings the same weights down to ~35 GB. The parameter count is what makes that arithmetic go.
- Inference cost. As a rule of thumb, each generated token costs roughly
2NFLOPs on a dense model withNparameters. At 70B, that’s ~140 billion FLOPs per token. This is why the same prompt is dramatically faster on an 8B model than a 70B one — and why “active” parameters matter so much for mixture-of-experts models. - Capability, loosely. Treat this one as a heuristic, not a law. Within a single model family trained the same way, the bigger sibling generally scores better — that’s what a family’s own model card usually shows. Across families it falls apart fast: a well-trained 8B beats a poorly trained 30B, and training data and post-training move results as much as size does. Use the parameter count for ballpark expectations; don’t use it to rank two well-engineered models from different labs.
The short answer
parameter = one learned number inside the model
Picture to keep: a stack of 80 spreadsheets of numbers, each about 8192 columns wide, and every generated token has to be dragged past every cell. The parameter count is how many cells there are. Except that the cells aren’t read in isolation — they’re multiplied in blocks, and in a mixture-of-experts model most of the stack is skipped for any given token, which is exactly where the headline number stops meaning what you think.
A parameter count is the total number of values (weights and biases) inside the neural network that were set by training rather than by a human writing code. “Llama 3.1 70B” means there are about 70 billion such numbers in the model file. Most of them live in matrices used by the transformer’s feed-forward and attention layers; multiplying input vectors through those matrices is most of what running the model means.
How it works
The interesting part is watching the simplest version of “how big is this model?” break, three times.
Naive attempt: just look at the file size. The 70B download is ~140 GB; the 8B one is ~16 GB. Done? Why it breaks: re-download the same 70B model quantized to 4 bits and the file is ~35 GB. The architecture didn’t change and neither did the number of learned values — only the precision each one is stored at (they’re rounded, so it isn’t free). File size mostly measures storage format. The invariant underneath is the count of learned numbers, and that’s the parameter count.
Where the parameters live
A transformer LLM is mostly stacks of
matrix multiplications.
For a model with hidden dimension
d and L layers, each layer contains roughly:
- Attention projections — four matrices (Q, K, V, output) each ~
d × din vanilla multi-head attention. ≈4d²parameters per layer. Modern variants like grouped-query attention shrink the K and V projections, so the real number is a bit less. - Feed-forward / MLP — two matrices, with hidden dim typically
~4d, giving~8d²parameters per layer. SwiGLU and similar variants use three matrices instead of two; the constant changes but the order is the same. - Layer norms and biases —
O(d)per layer. Basically rounding error.
Adding up: a “vanilla” transformer layer is roughly 12d² parameters. With
L layers, the body of the model is ~12 · L · d². The
embedding and output tables
add ~2 · V · d where V is the vocab size; for big vocabularies
(100k+) this is a couple of percent of the total.
For Llama 3.1 70B
(d = 8192, L = 80, V ≈ 128k), the back-of-envelope
is 12 × 80 × 8192² ≈ 64B from the body plus a couple of billion for
embeddings — not exactly 70B, but close enough that the math is recognizably
doing the right thing. The remaining gap comes from architecture specifics
(SwiGLU’s three matrices, the precise FFN width, untied output embeddings).
The full architecture table is in the Llama 3 paper; the formula above is
the part worth carrying in your head.
The takeaway: the FFN matrices hold most of the parameters — typically around two-thirds of the body. This is why the FFN is what mixture-of-experts replaces. That’s where the budget is.
What the parameters cost
So now you can count them. Second naive attempt: assume the count tells you what the model costs to run. For a dense model this mostly holds, and it’s worth seeing exactly why. Two physical costs scale linearly with parameter count:
- Storage / memory — bytes per param × N. In bf16, 2 bytes. In fp8, 1 byte. In int4, 0.5 bytes. The model file and the VRAM to hold the weights are basically N × bytes-per-param, plus framework overhead and (for quantized formats) the per-block scale factors. During training you also pay for optimizer state, which scales with N: Adam keeps two extra numbers per parameter, and mixed-precision setups typically keep a full-precision master copy of the weights on top of that. The KV cache and activations scale with batch and sequence length and the model’s hidden dim — not directly with N.
- FLOPs per token — roughly
2Nfor a forward pass on a dense model. Each weight participates in one multiply and one add per token of input. The backward pass during training is about another4N, which is where the rule of thumb “training a model on D tokens costs ≈6NDFLOPs” (Kaplan et al., 2020) comes from.
The 2N-per-token relationship is why inference cost is so legible. Doubling
the parameter count roughly doubles the per-token cost. It’s also why
mixture-of-experts is interesting: an MoE model only fires a subset of its
parameters per token, so total params and active params come apart.
What gets counted
Why that breaks: the assumption hiding inside “N parameters ⇒ 2N FLOPs per token” is that every parameter is used for every token. Once that stops being true, the headline number stops being a cost estimate — and a few other accounting choices quietly move it too.
- Embeddings. Some early scaling-law work (Kaplan et al., 2020) reports non-embedding parameter counts, on the grounds that embeddings don’t scale the same way as the rest of the network. Modern model cards generally report total parameters, embeddings included — but it’s worth checking rather than assuming, because the two numbers can differ by a few percent, and by much more for small models with big vocabularies.
- Tied vs. untied embeddings. Some models share weights between the input embedding and the output projection (tied); some don’t (untied). Tied counts those parameters once, untied counts them twice. Two architecturally similar models can land at different totals just from this.
- Active vs. total params (MoE). A mixture-of-experts model has many expert FFNs but routes only a couple per token. DeepSeek-V3 has 671B total parameters but only ~37B active per token. “DeepSeek-V3 671B” looks like a 671B model on the shelf and costs more like a 37B model to run per token. The headline number stops being a single-axis comparison the moment MoE enters.
- Quantization. A 70B model run in int4 still has 70B parameters — the count doesn’t change — but each parameter takes fewer bytes (and may be slightly less accurate). Quantization changes bytes-per-param, not N.
So when you read “X B parameters,” ask: is it dense or MoE? (If MoE: total or active?) Total or non-embedding? For most dense model cards the headline is total parameters unless stated otherwise, and the math above lines up.
You started with parameter = one learned number inside the model. What did
the three breakages add? — + N is the storage-format-independent invariant,
+ cost ≈ 2N FLOPs/token only when every parameter fires, and + the number on the box is an accounting convention, not a measurement. That’s why
“DeepSeek-V3 671B” and “Llama 3.1 70B” are only comparable on one axis: both
tell you shelf size, and only the dense one also tells you roughly what a
token costs.
Famous related terms
- Weights —
weights = the learned matrix entries— usually used loosely as a synonym for “parameters.” Strictly, the parameter total is weights plus biases plus embedding tables; “download the weights” means all of it. - Active parameters (MoE) —
active params = params actually used per token— the relevant cost number for mixture-of-experts models, often much smaller than the total. - FLOPs per token —
FLOPs/token ≈ 2N(dense, forward pass) — the rule of thumb that turns parameter count into inference cost. - bf16 / fp8 / int4 —
bytes per param = how the parameter is stored— quantization shrinks bytes per param without changing the parameter count itself. - Scaling laws —
loss falls as a power law in (N, D, C)— the empirical fact that made parameter count the headline number in the first place. See why scaling laws exist. - VRAM —
VRAM needed ≈ bytes/param × N + KV cache + activations— why parameter count is the first thing checked when GPU memory is the bottleneck. - Mixture of experts —
MoE = many FFN experts + a router— decouples total params from active params. The model can be huge on disk and cheap per token at the same time.
Going deeper
- The Llama 3 Herd of Models (Meta, 2024) — the primary source, and the way to settle “where exactly do the missing 6 billion parameters come from?”: its architecture tables give the real
d,L, FFN width and vocab for 8B/70B/405B. - Andrej Karpathy’s Let’s build GPT — answers “what does
12·L·d²feel like?”, by building a tiny transformer and watching the count grow as you change one hyperparameter. - Attention Is All You Need (Vaswani et al., 2017) — the rabbit hole, for “why are there four attention matrices and not one?”