What are 'weights' in an LLM?
When Meta releases 'open weights' for Llama, what's actually in that file? A giant table of numbers and nothing else — so how does a pile of numbers know things?
On this page
The picture version
Five pictures for a reader who has never opened one of these files. The prose below fills in the seams the pictures skip.
1 · The problem
140 gigabytes arrive. None of it is instructions.
2 · What one of them is
A single number at a fixed address, doing one multiplication.
3 · The obvious guess, and its failure
There is no row that says France → Paris.
4 · What is literally in the file
A length, a table of contents, and a wall of bytes.
5 · Keep this card
The whole thing on one index card.
Why it exists
You go to Hugging Face, click Download on Llama 3.1 70B, watch your disk
space evaporate for an hour, and end up with a folder of
.safetensors
files totalling roughly 140 gigabytes (71 billion parameters at 2 bytes
each). Alongside them sit a config file and some tokenizer assets measured in
kilobytes. And that’s the whole model. No source code that says “when asked
about Paris, mention the Eiffel Tower.” No database of facts. No rules. Just
an enormous table of numbers.
Those numbers are the weights. When people say “Meta released the weights” or “the model is 140 GB” or “open-weights model,” they’re talking about those files. The architecture — how the layers connect — is a modest amount of code that’s public and re-implementable. The weights are what makes the running program Llama rather than a random untrained network outputting gibberish.
The word survives from the original mental picture of a neural network: each connection between artificial neurons has a strength — a weight — that says how much one neuron’s output feeds into the next. Adjust the weights and you change the function the network computes. Train on enough text and the weights settle into a configuration that, when you feed in tokens and multiply them through, produces output that looks like a fluent answer.
Why it matters now
Three places the weight file shows up as more than a definition:
- “Open weights” vs “open source.” Meta ships Llama under a license it calls open source, but the OSI’s Open Source AI Definition asks for three things — the parameters, the code, and detailed data information about what the system was trained on. A weights-only release satisfies the first and not the last, which is my read of why Llama-style licenses sit outside that definition. You can run, fine-tune, and quantize the released weights — but you can’t reproduce them from scratch. The dispute over what “open” should mean in AI falls almost exactly along the line between releasing the weights and releasing what made them.
- What gets backed up, copied, leaked. A model release is a weight release. Shipping a copy of the file is, functionally, shipping the model — which is enough to explain, without needing anyone’s internal policy, why labs that don’t release weights guard those files closely.
- What fine-tuning, quantization, distillation, and merging actually do. Fine-tuning, quantization, and merging edit or re-encode an existing weight file; distillation trains a new, smaller one to imitate the old one’s outputs. Holding “the weight file is a big array of numbers in named matrices” in your head is what makes the rest of the toolkit legible.
The short answer
weights = the numbers inside the model's matrices, set by training
Picture to keep: not a filing cabinet you look things up in, but the fixed shape of a pinball machine — millions of pegs, each one nudging whatever passes through it, and the answer is where the ball comes out.
A neural network is mostly matrix multiplications. The weights are the entries in those matrices. They’re set during training by gradient descent and then frozen. Running the model means pushing token vectors through those matrices: multiply, add, repeat. Everything the model learned is encoded in those numbers and nowhere else — the rest of what you downloaded is plumbing that decides how they get used.
How it works
The natural first guess about that 140 GB file is that it’s a database: somewhere in there, a record that pairs “France” with “Paris.” Chase that guess and watch it fail — the failure is where the actual mechanism lives.
First, what a single weight is
A weight is one floating-point number. In current open-weights releases it’s typically 16 bits (bf16), sometimes quantized down to 8 or 4 bits for inference.
It lives at a fixed position in a specific matrix in a specific layer. Say
row 4096, column 2071, of the down-projection matrix in layer 42’s
feed-forward block. The number might be -0.00347. By itself, that number
means nothing.
What it does is mechanical: when an input vector passes through that
matrix, the value at column 2071 of the input gets multiplied by -0.00347
and added into row 4096 of the output. That tiny contribution combines with
millions of others to produce the next layer’s input. There is no “this
weight means cat” — every weight participates in millions of dot products,
and every output is a weighted sum of all of them.
So where’s the database?
There isn’t one. That’s the failure of the first guess, and it’s worth sitting with, because it feels suspicious: no lookup table, no row that says “France → Paris.” Training adjusts weights until the act of running them — multiplying a sequence of vectors through every layer — produces output probabilities that match the training distribution.
The result is that knowledge ends up distributed. A fact like “Paris is the capital of France” isn’t stored in one weight or one neuron. It’s spread across many weights in many layers, all of which also participate in encoding millions of other facts. Zeroing out one weight out of tens of billions is not expected to change much of anything; degrade enough of them and behavior falls apart, but rarely in a clean “it forgot France” way.
The closest thing anyone can currently point to for “where a concept lives” is the attention heads and circuits that activate when that concept shows up. Mechanistic interpretability is the research program trying to reverse-engineer those circuits — naming subsets of weights that, together, implement something a human can describe. Progress is real but partial. The honest summary in 2026: for any specific weight in a frontier-size model, we generally cannot say what it does in isolation.
This is also why the pinball picture beats the filing-cabinet one — but note where it breaks: a real pinball machine is chaotic, and these pegs are tuned so that similar inputs land in nearly the same place. The point of the analogy is only that the knowledge is in the arrangement, not in any peg.
Weights vs parameters
The two words get used interchangeably and most of the time that’s fine. The technical distinction:
- Weights are the entries of the matrices — the multipliers in
output = W · input + b. - Biases are the per-row additive offsets — the
bin the same equation. - Parameters = weights + biases + any other learned scalars (e.g. layer-norm scales).
Biases and norm scales are a small fraction of the total, because they scale
with hidden dimension d while weight matrices scale with d² — and several
current model families drop the linear biases entirely. So the “X B parameters”
headline on a model card is dominated by weights, and “weight count” and
“parameter count” come out close enough that people use them
interchangeably. The headline number is the count; the weights is the
contents.
What’s literally in the file
A .safetensors file (the modern standard format) is three parts laid
end to end:
- An 8-byte little-endian integer giving the length of the JSON header that follows.
- A JSON header with one entry per tensor — name, shape, dtype,
and a
data_offsetspair giving start and end positions within the byte buffer (not absolute file offsets). E.g.model.layers.42.mlp.down_proj.weight, shape[8192, 28672], dtypebf16. - The raw tensor bytes, packed contiguously.
That’s the learned state of the model. You still need the architecture code
and the tokenizer files (vocabulary + merge rules) to
turn the files into something that takes a prompt and emits text — but the
learned state, the part months of training produced, is exactly the
.safetensors bytes.
You started with weights = the numbers inside the model's matrices, set by training. What did this post add? — + and nowhere else. Everything the
model appears to know lives in that one arrangement of numbers, which is why
downloading the file is downloading the model, why fine-tuning and
quantization and merging are all just different ways of editing it, and why
“which weight knows about France?” has no answer.
Famous related terms
- Parameters —
parameters = weights + biases + other learned scalars— the count of all trainable numbers. See what ‘X parameters’ means for why that number is the headline on every model card. - Open-weights model —
open-weights = the trained weight file is downloadable— distinct from open-source under OSI’s definition, which also asks for the code and the data information behind the weights. - Checkpoint —
checkpoint = a snapshot of the weights at one point in training— saved periodically during a long run so it can resume from the last snapshot instead of from zero when a node dies. - Gradient descent —
gradient descent = nudge each weight against the slope of the loss— the algorithm that produces the weights from training data. - Fine-tuning —
fine-tuning = continue training a model's weights on new data— modifies the same weight file rather than starting from scratch. See why fine-tuning is cheap. - Distillation —
distillation = train a small model to copy a big model's outputs— produces a smaller weight file that mimics a larger one. See why distillation exists. - Quantization —
quantization = store each weight in fewer bits— doesn’t change which weights exist, only how they’re encoded. See why quantization works.
Going deeper
- The safetensors format spec —
the primary source for what’s literally in the file you downloaded: the
header-length prefix, the JSON header with
data_offsets, then the raw bytes. - Andrej Karpathy’s Neural Networks: Zero to Hero — the best explainer for what a weight does: builds a network from scratch and shows what gradient descent does to it step by step, until “a weight is just a number in a matrix” stops feeling abstract.
- Anthropic’s A Mathematical Framework for Transformer Circuits — the rabbit hole for “where does knowledge live in those numbers”: the entry point into mechanistic interpretability and treating weights as circuits you can reverse-engineer.