Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why LoRA exists

Full fine-tuning a 70B model means storing optimizer state for 70 billion weights. LoRA trains under 1% of the parameters and, on the tasks people have tested, often matches the result. The trick is a hypothesis about the shape of the update.

AI & ML intermediate Apr 29, 2026 · updated Aug 25, 2026 · 15 min read

On this page

The picture version

Six pictures for a reader who has never fine-tuned anything. The prose below fills in the seams the pictures skip.

1 · The problem

Five customers. Five near-identical 70B files.

base model 70B fine-tune ×5 70B70B70B70B70B weight for weight, almost identical files and each run needed ~560 GB of Adam optimizer state alone You don’t need 70 billion degrees of freedom to say “talk like this customer”.
Adam keeps two extra 32-bit numbers per parameter — a running mean and variance of the gradients — on top of the weights and the gradients themselves. One base model, five customers is the running example for this post, and both the disk bill and the memory bill come from treating each fine-tune as a whole new model.

2 · The hypothesis

The update is small, even though the model isn’t.

the clue (Aghajanyan et al., 2020) fine-tune RoBERTa to ~90% of full performance on MRPC by optimising ~200 parameters the next step (Hu et al., 2021) don’t sample a random subspace — learn one, by baking the constraint into the parameterisation itself so the bet is: ΔW ≈ B A, with r much smaller than either side what the evidence shows: if you constrain the update to be low-rank, you don’t lose much. it does not show that a full fine-tune only ever moves weights in a low-rank subspace — the deeper claim is still under active debate
The technique rests on a guess that turned out right in practice: the adaptation a fine-tune performs lives in a low-dimensional subspace even though the model is huge. Operationally that is all you needed — and it is worth keeping the weaker claim and the stronger one apart.

3 · The mechanics

One fat frozen matrix, two skinny trainable ones.

W₀ 4096 × 4096 frozen never moves + B 4096 × 8 A 8 × 4096 the only place rank can be lost ~16.8M parameters per matrix, before ~65k trainable, after — about 256× fewer And B starts at zero, so the adapted model begins as an exact copy of the base.
A rank-r update is a sum of r outer products; BA is just that, factored. Training touches only A and B — and at serving time you can fold them back into the weights and pay no inference penalty at all.

4 · Where the saving actually comes from

Not the parameters. The optimizer state hanging off them.

full fine-tune — what has to live in memory weights gradients optimizer state — two FP32 numbers per parameter with LoRA — the optimizer state attaches only to the trainable slice weights (frozen) that’s the whole trainable footprint adapters are tiny on disk tens to a few hundred MB, against ~14 GB for a 7B base at FP16 and stacking quantization goes further QLoRA: 4-bit frozen base + 16-bit adapters, 65B fine-tuned on a single 48 GB GPU Your five customers now need one machine, not one cluster.
Cutting trainable parameters by roughly a thousandfold cuts the thing that was actually blowing up VRAM. The small adapter files are the pleasant side effect; the memory is the reason it works on hardware you own. Bar widths are illustrative.

5 · What it isn’t

Parameter-efficient. Not knowledge-efficient.

what it’s for tone and style output format task framing, refusal boundaries the adaptation the low-rank bet covers best what it’s bad at stuffing in new facts anything that changes weekly making up for thin training data retrieval handles all three — a database write beats a retrain It shrinks how many weights update, not how much data you need. if your fine-tune is bad, more rank won’t save you — better data will
One more boundary worth holding: LoRA and a full fine-tune reach different solutions in weight space even when the downstream metrics match, with different generalisation and forgetting behaviour. Whether they are equivalent in any deeper sense is an open dispute, so evaluate against a full fine-tune when the task is hard or the data is large.

6 · Keep this card

The whole thing on one index card.

LoRA = freeze the base weights + add a low-rank update ΔW = BA + train only A and B ∴ the saving is the optimizer state, not the parameter count
Picture to keep: the base model is a printed textbook you’re not allowed to write in, and a LoRA is a stack of transparent overlays you clip on top — each customer gets their own sheet, the book never changes, and swapping customers means swapping a sheet. Where it breaks: the overlay isn’t a margin note. It is added to the page’s contents before you read them, so it can change what the book says.

Why it exists

If you’ve ever downloaded a small file from an image-model community and dropped it next to a multi-gigabyte model to make it draw in one specific style, you’ve already used the thing this post is about. The file was somewhere between tens and a few hundred megabytes. The model it modified was orders of magnitude bigger and didn’t change at all. That’s the trick — and here’s the problem it was invented to solve.

Your team fine-tunes a 70-billion-parameter model for one customer, using a few thousand examples of how they want it to behave. It works. A second customer asks for the same thing on their own data. Then a third. Six weeks later you are paying to store five separate 70B checkpoints that are, weight for weight, almost identical files — and each one took a GPU cluster and a mountain of optimizer state to produce. One base model, five customers: that’s the running example for this post.

The naive plan is the obvious one — load the model, run gradient descent on every weight, save the result — and it’s brutal in both directions. Adam, the optimizer everyone reaches for, keeps two extra FP32 numbers per parameter (the running mean and variance of gradients). On a 70B model, that’s ~560 GB of optimizer state alone, on top of the weights and gradients. And then there’s the disk bill from the hook: five customers means five full copies of a 70B-parameter model.

LoRA — short for Low-Rank Adaptation — exists because someone noticed the whole setup was wasteful. The update you’re trying to learn during fine-tuning — the difference between “base model” and “base model that does my task” — turns out to have a very particular shape. You don’t need 70 billion degrees of freedom to express it. A few million is usually enough.

That’s the entire pitch. Freeze the base. Train a tiny side-channel. Match the quality of full fine-tuning at a fraction of the cost.

Why it matters now

LoRA isn’t a clever optimization buried in some lab’s training code. It’s the default first thing people reach for under the PEFT umbrella, and it shapes the infrastructure around it.

If you’re choosing between fine-tuning, prompting, and RAG, LoRA is what changes the price of the fine-tuning option — often from “we can’t” to “we can, per customer.” It doesn’t collapse the choice (prompting and retrieval solve different problems), but it moves where the line sits. Knowing why LoRA works tells you when it’s the right tool.

The short answer

LoRA = freeze base weights + add a low-rank update ΔW = BA + only train A and B

Picture to keep: the base model is a printed textbook you’re not allowed to write in. A LoRA is a stack of transparent overlays you clip on top — each customer gets their own sheet, the book underneath never changes, and swapping customers means swapping a sheet, not reprinting the book. Like that, except the overlay isn’t a note in the margin: it’s added to every relevant page’s contents before you read them, so it can change what the book says, not just annotate it.

Instead of learning a new full-size weight matrix, you learn two skinny matrices whose product approximates the update. If W is 4096×4096 and your rank r is 8, then B is 4096×8 and A is 8×4096. That’s ~65k trainable numbers in place of ~16.8M — about 256× fewer parameters, matching most of the quality.

How it works

The design is a chain of “that breaks, so do this instead,” and it starts with a question nobody asked for a while: does the update actually need to be full-size?

The intrinsic-rank hypothesis

The whole technique rests on a guess that turned out to be right in practice: the adaptation a fine-tune performs lives in a low-dimensional subspace, even though the model itself is huge.

This wasn’t pulled out of nowhere. Aghajanyan et al. (2020), in Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning, showed you can fine-tune RoBERTa to ~90% of full performance on MRPC by optimizing only ~200 parameters projected randomly back into the full weight space. One reading of that result: the pretrained model already has the right features, and fine-tuning just steers them. Hu et al. (2021) took the next step: if the effective update is low-dimensional, why not bake that constraint into the parameterization itself? Don’t sample a random subspace — learn one, by writing the update as a product of two low-rank matrices.

So the hypothesis is: ΔW (the change to a weight matrix during fine-tuning) is well-approximated by BA, where B ∈ ℝ^{d×r} and A ∈ ℝ^{r×k} and r is much smaller than d or k.

The empirical result: yes, often it is. r=8 frequently matches r=64 on downstream tasks. The QLoRA paper, in an Alpaca-style sweep on LLaMA-7B with LoRA applied to all layers, reported that rank had effectively no effect on final task performance — a narrower observation than “rank doesn’t matter,” but a striking one.

What this doesn’t prove: it doesn’t show that a full fine-tune “actually” only moves weights in a low-rank subspace. It only shows that if you constrain the update to be low-rank, you don’t lose much. Which is what you wanted operationally; the deeper claim is still under active debate.

The mechanics

For a frozen pretrained weight matrix W₀ ∈ ℝ^{d×k}, the LoRA parameterization replaces

y = W₀ x

with

y = W₀ x + (α/r) · B A x

where B ∈ ℝ^{d×r}, A ∈ ℝ^{r×k}, and α is a scalar scaling factor. In the original paper’s scheme, B is initialized to zero and A to random Gaussian, so BA = 0 at step 0 and the adapted model starts as an exact copy of the base. (Library defaults differ in details — e.g. Hugging Face PEFT uses Kaiming-uniform for A and zeros for B — but the “starts at zero update” property is preserved.) Training updates only A and B; W₀ never moves.

A few details worth knowing:

Why the savings are this dramatic

Stack a few effects:

QLoRA: stacking the trick on quantization

QLoRA (Dettmers et al., May 2023) is the most consequential follow-up. The base model is quantized to 4-bit (NF4, a data type tailored for normally-distributed weights), kept frozen, and dequantized on the fly to 16-bit (FP16/BF16) to compute forward and backward passes. Only the LoRA adapters and the optimizer state for them live in 16-bit. The headline: fine-tune a 65B model on one 48 GB GPU, recovering full 16-bit fine-tuning quality on the benchmarks they tested. A 48 GB card is still a serious purchase, but it’s a single card — which is what moved 65B fine-tuning out of cluster territory. Your five customers now need one machine, not one cluster.

Where it gets subtle

The whole technique is a bet on a structural observation about fine-tuning, not a clever optimizer trick. That’s why it generalizes — the same idea works for transformer LLMs, diffusion image models, and just about anything else where you start from a strong pretrained base. Your five customers now cost one base model plus five files small enough to sit side by side on a single server.

You started with LoRA = freeze base weights + a low-rank update BA. What did this post add that the line hides? — + the savings come from the optimizer state, not the parameter count. Adam keeps two extra FP32 numbers per trainable parameter, so cutting trainable parameters by ~1000× cuts the thing that was actually blowing up your VRAM. The tiny adapter files are the pleasant side effect; the memory is the reason it works on hardware you own.

Check yourself

Before you go — someone reports that bumping LoRA rank from 8 to 128 barely changed their eval scores, and concludes “rank doesn’t matter, always use 8.” What’s wrong with that inference?

Answer

They’ve generalized from one task. The flat rank-vs-quality curve is a real and widely reported observation — QLoRA’s Alpaca-style sweep on LLaMA-7B found rank had effectively no effect on final task performance — but it’s an observation about the tasks people have tested, mostly style-and-format adaptation on modest datasets. The intrinsic-rank hypothesis says the update for that task lives in a low-dimensional subspace; it doesn’t say every conceivable adaptation does. If their fine-tune is small and stylistic, r=8 is a fine default. If it’s large and genuinely teaches new behavior, the honest move is to sweep rank and also compare against a full fine-tune — which is exactly what the Illusion of Equivalence line of work argues for.

Next — a sixth customer arrives. They want the model to answer in formal Japanese, refuse anything outside their product line, and always end with a specific disclaimer. They have 4,000 example conversations. Is this a LoRA job, and what would you predict about rank?

Answer

Yes — this is close to the ideal case. Everything they’re asking for is style, format and task framing: tone, refusal boundary, output template. That’s the kind of adaptation the intrinsic-rank hypothesis covers best, and it’s the kind that has produced the famously flat rank-vs-quality curve. Start at r=8, expect it to be enough, and spend your effort on the 4,000 examples rather than on the rank sweep — LoRA is parameter-efficient, not data-efficient, so the data is the thing that decides whether it works. Ship the adapter alongside the other five; the base model doesn’t change.

And one more — a colleague wants to use LoRA to teach a model your company’s internal product catalog, which changes weekly. Good plan?

Answer

Probably not, for two separate reasons. First, LoRA is parameter-efficient, not knowledge-efficient — it shrinks how many weights update, not how much data you need, and fine-tuning in general is a poor tool for injecting specific facts compared to retrieval. Second, “changes weekly” means you’d be retraining weekly and the model would carry stale facts between runs with no way to invalidate them. Retrieval handles both: put the catalog in a store, fetch the relevant rows at query time, and updating it is a database write. Save the LoRA for teaching the model how to talk about the catalog.

Going deeper