Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

What are 'weights' in an LLM?

When Meta releases 'open weights' for Llama, what's actually in that file? A giant table of numbers and nothing else — so how does a pile of numbers know things?

AI & ML intro May 16, 2026 · updated Aug 25, 2026 · 9 min read

On this page

The picture version

Five pictures for a reader who has never opened one of these files. The prose below fills in the seams the pictures skip.

1 · The problem

140 gigabytes arrive. None of it is instructions.

Download · Llama 3.1 70B an hour later .safetensors ≈ 140 GB 71 billion numbers, 2 bytes each config + tokenizer kilobytes what is not in there code saying “for Paris, mention the Eiffel Tower” a database of facts rules of any kind Just an enormous table of numbers.
Those numbers are the weights, and they are the whole model — the architecture is a modest amount of public code, and the tokenizer assets are kilobytes. So the question the rest of these pictures answer is: how does a pile of numbers know anything?

2 · What one of them is

A single number at a fixed address, doing one multiplication.

its address: layer 42 → feed-forward block → down-projection matrix → row 4096, column 2071 -0.00347 one weight, one floating-point number, 16 bits by itself it means nothing at all input, at col 2071 × -0.00347 added into output, row 4096 one tiny contribution, combined with millions of others to make the next layer’s input There is no “this weight means cat.”
Every weight participates in millions of dot products, and every output is a weighted sum of all of them. A weight has a job, not a meaning — which is exactly why the next scene’s obvious guess fails.

3 · The obvious guess, and its failure

There is no row that says France → Paris.

the guess: a filing cabinet France → Paris Japan → Tokyo Peru   → Lima no lookup table. no such row. what is actually there millions of pegs, each nudging what passes through Knowledge is in the arrangement, not in any one number. zero out one weight in tens of billions and nothing much happens — degrade enough of them and behaviour falls apart, but rarely in a clean “it forgot France” way
“Paris is the capital of France” is spread across many weights in many layers, all of which also help encode millions of other facts. Where the pinball picture breaks: a real machine is chaotic, while these pegs are tuned so similar inputs land in nearly the same place.

4 · What is literally in the file

A length, a table of contents, and a wall of bytes.

8B a little-endian integer: how long the header is JSON header — one entry per tensor model.layers.42.mlp.down_proj.weight shape [8192, 28672] · dtype bf16 data_offsets: start, end names, shapes, types, and where each block sits inside the byte buffer 3f 80 be 41 c2 07 3d 91 bf 12 40 aa 3e 6c c1 08 be 44 40 1a be 77 3f 05 c1 6e 3e 92 bf 30 40 08 3d d4 c2 11 bf 41 3e 09 40 6d c0 88 3f 24 be 5b 41 03 3d 96 bf 71 c1 55 3e 8c 3f 60 bf 19 40 47 be 21 3d af c2 30 3f 08 the raw tensor bytes, packed contiguously — this is the part that took months of training to produce a few kilobytes of bookkeeping essentially all 140 GB The header tells you where things are. The bytes are what was learned. bytes shown are made up for the picture — real ones are not readable in any useful sense
You still need the architecture code and the tokenizer files to turn this into something that takes a prompt and emits text. But the learned state — the part training actually produced — is exactly these bytes, which is why shipping the file is functionally shipping the model.

5 · Keep this card

The whole thing on one index card.

weights = the numbers inside the model’s matrices + set by training, then frozen + and nowhere else
Picture to keep: not a filing cabinet you look things up in, but the fixed shape of a pinball machine — millions of pegs, each nudging whatever passes through, and the answer is where the ball comes out. That is why downloading the file is downloading the model, and why “which weight knows about France?” has no answer.

Why it exists

You go to Hugging Face, click Download on Llama 3.1 70B, watch your disk space evaporate for an hour, and end up with a folder of .safetensors files totalling roughly 140 gigabytes (71 billion parameters at 2 bytes each). Alongside them sit a config file and some tokenizer assets measured in kilobytes. And that’s the whole model. No source code that says “when asked about Paris, mention the Eiffel Tower.” No database of facts. No rules. Just an enormous table of numbers.

Those numbers are the weights. When people say “Meta released the weights” or “the model is 140 GB” or “open-weights model,” they’re talking about those files. The architecture — how the layers connect — is a modest amount of code that’s public and re-implementable. The weights are what makes the running program Llama rather than a random untrained network outputting gibberish.

The word survives from the original mental picture of a neural network: each connection between artificial neurons has a strength — a weight — that says how much one neuron’s output feeds into the next. Adjust the weights and you change the function the network computes. Train on enough text and the weights settle into a configuration that, when you feed in tokens and multiply them through, produces output that looks like a fluent answer.

Why it matters now

Three places the weight file shows up as more than a definition:

The short answer

weights = the numbers inside the model's matrices, set by training

Picture to keep: not a filing cabinet you look things up in, but the fixed shape of a pinball machine — millions of pegs, each one nudging whatever passes through it, and the answer is where the ball comes out.

A neural network is mostly matrix multiplications. The weights are the entries in those matrices. They’re set during training by gradient descent and then frozen. Running the model means pushing token vectors through those matrices: multiply, add, repeat. Everything the model learned is encoded in those numbers and nowhere else — the rest of what you downloaded is plumbing that decides how they get used.

How it works

The natural first guess about that 140 GB file is that it’s a database: somewhere in there, a record that pairs “France” with “Paris.” Chase that guess and watch it fail — the failure is where the actual mechanism lives.

First, what a single weight is

A weight is one floating-point number. In current open-weights releases it’s typically 16 bits (bf16), sometimes quantized down to 8 or 4 bits for inference.

It lives at a fixed position in a specific matrix in a specific layer. Say row 4096, column 2071, of the down-projection matrix in layer 42’s feed-forward block. The number might be -0.00347. By itself, that number means nothing.

What it does is mechanical: when an input vector passes through that matrix, the value at column 2071 of the input gets multiplied by -0.00347 and added into row 4096 of the output. That tiny contribution combines with millions of others to produce the next layer’s input. There is no “this weight means cat” — every weight participates in millions of dot products, and every output is a weighted sum of all of them.

So where’s the database?

There isn’t one. That’s the failure of the first guess, and it’s worth sitting with, because it feels suspicious: no lookup table, no row that says “France → Paris.” Training adjusts weights until the act of running them — multiplying a sequence of vectors through every layer — produces output probabilities that match the training distribution.

The result is that knowledge ends up distributed. A fact like “Paris is the capital of France” isn’t stored in one weight or one neuron. It’s spread across many weights in many layers, all of which also participate in encoding millions of other facts. Zeroing out one weight out of tens of billions is not expected to change much of anything; degrade enough of them and behavior falls apart, but rarely in a clean “it forgot France” way.

The closest thing anyone can currently point to for “where a concept lives” is the attention heads and circuits that activate when that concept shows up. Mechanistic interpretability is the research program trying to reverse-engineer those circuits — naming subsets of weights that, together, implement something a human can describe. Progress is real but partial. The honest summary in 2026: for any specific weight in a frontier-size model, we generally cannot say what it does in isolation.

This is also why the pinball picture beats the filing-cabinet one — but note where it breaks: a real pinball machine is chaotic, and these pegs are tuned so that similar inputs land in nearly the same place. The point of the analogy is only that the knowledge is in the arrangement, not in any peg.

Weights vs parameters

The two words get used interchangeably and most of the time that’s fine. The technical distinction:

Biases and norm scales are a small fraction of the total, because they scale with hidden dimension d while weight matrices scale with d² — and several current model families drop the linear biases entirely. So the “X B parameters” headline on a model card is dominated by weights, and “weight count” and “parameter count” come out close enough that people use them interchangeably. The headline number is the count; the weights is the contents.

What’s literally in the file

A .safetensors file (the modern standard format) is three parts laid end to end:

That’s the learned state of the model. You still need the architecture code and the tokenizer files (vocabulary + merge rules) to turn the files into something that takes a prompt and emits text — but the learned state, the part months of training produced, is exactly the .safetensors bytes.

You started with weights = the numbers inside the model's matrices, set by training. What did this post add? — + and nowhere else. Everything the model appears to know lives in that one arrangement of numbers, which is why downloading the file is downloading the model, why fine-tuning and quantization and merging are all just different ways of editing it, and why “which weight knows about France?” has no answer.

Going deeper