Why LoRA exists
Full fine-tuning a 70B model means storing optimizer state for 70 billion weights. LoRA trains under 1% of the parameters and, on the tasks people have tested, often matches the result. The trick is a hypothesis about the shape of the update.
On this page
The picture version
Six pictures for a reader who has never fine-tuned anything. The prose below fills in the seams the pictures skip.
1 · The problem
Five customers. Five near-identical 70B files.
2 · The hypothesis
The update is small, even though the model isn’t.
3 · The mechanics
One fat frozen matrix, two skinny trainable ones.
4 · Where the saving actually comes from
Not the parameters. The optimizer state hanging off them.
5 · What it isn’t
Parameter-efficient. Not knowledge-efficient.
6 · Keep this card
The whole thing on one index card.
Why it exists
If you’ve ever downloaded a small file from an image-model community and dropped it next to a multi-gigabyte model to make it draw in one specific style, you’ve already used the thing this post is about. The file was somewhere between tens and a few hundred megabytes. The model it modified was orders of magnitude bigger and didn’t change at all. That’s the trick — and here’s the problem it was invented to solve.
Your team fine-tunes a 70-billion-parameter model for one customer, using a few thousand examples of how they want it to behave. It works. A second customer asks for the same thing on their own data. Then a third. Six weeks later you are paying to store five separate 70B checkpoints that are, weight for weight, almost identical files — and each one took a GPU cluster and a mountain of optimizer state to produce. One base model, five customers: that’s the running example for this post.
The naive plan is the obvious one — load the model, run gradient descent on every weight, save the result — and it’s brutal in both directions. Adam, the optimizer everyone reaches for, keeps two extra FP32 numbers per parameter (the running mean and variance of gradients). On a 70B model, that’s ~560 GB of optimizer state alone, on top of the weights and gradients. And then there’s the disk bill from the hook: five customers means five full copies of a 70B-parameter model.
LoRA — short for Low-Rank Adaptation — exists because someone noticed the whole setup was wasteful. The update you’re trying to learn during fine-tuning — the difference between “base model” and “base model that does my task” — turns out to have a very particular shape. You don’t need 70 billion degrees of freedom to express it. A few million is usually enough.
That’s the entire pitch. Freeze the base. Train a tiny side-channel. Match the quality of full fine-tuning at a fraction of the cost.
Why it matters now
LoRA isn’t a clever optimization buried in some lab’s training code. It’s the default first thing people reach for under the PEFT umbrella, and it shapes the infrastructure around it.
- One base model, many adapters. A LoRA adapter for a 7B model is typically tens of megabytes. You can keep many of them on a single server and hot-swap between customers, languages, or styles without reloading the base weights. That’s what makes per-customer fine-tunes — the five customers from the hook — a serving problem rather than a capacity problem.
- Single-GPU fine-tuning. Combined with 4-bit quantization (the QLoRA recipe, Dettmers et al., 2023), the authors report fine-tuning a 65B-parameter model on a single 48 GB GPU.
- The open-weight ecosystem leans on this. Full fine-tunes of a 70B model are out of reach for hobbyists; a LoRA isn’t, which is why community fine-tunes are so often adapters. Nobody publishes a census of the hub, so read that as the shape of the incentive rather than a measured share.
- Image-model land too. The character and style packs traded in Stable Diffusion communities are usually LoRA adapters — that’s the small file from the hook. Same trick, different domain.
If you’re choosing between fine-tuning, prompting, and RAG, LoRA is what changes the price of the fine-tuning option — often from “we can’t” to “we can, per customer.” It doesn’t collapse the choice (prompting and retrieval solve different problems), but it moves where the line sits. Knowing why LoRA works tells you when it’s the right tool.
The short answer
LoRA = freeze base weights + add a low-rank update ΔW = BA + only train A and B
Picture to keep: the base model is a printed textbook you’re not allowed to write in. A LoRA is a stack of transparent overlays you clip on top — each customer gets their own sheet, the book underneath never changes, and swapping customers means swapping a sheet, not reprinting the book. Like that, except the overlay isn’t a note in the margin: it’s added to every relevant page’s contents before you read them, so it can change what the book says, not just annotate it.
Instead of learning a new full-size weight matrix, you learn two skinny
matrices whose product approximates the update. If W is 4096×4096 and
your rank r is 8, then B is 4096×8 and A is 8×4096. That’s ~65k
trainable numbers in place of ~16.8M — about 256× fewer parameters,
matching most of the quality.
How it works
The design is a chain of “that breaks, so do this instead,” and it starts with a question nobody asked for a while: does the update actually need to be full-size?
The intrinsic-rank hypothesis
The whole technique rests on a guess that turned out to be right in practice: the adaptation a fine-tune performs lives in a low-dimensional subspace, even though the model itself is huge.
This wasn’t pulled out of nowhere. Aghajanyan et al. (2020), in Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning, showed you can fine-tune RoBERTa to ~90% of full performance on MRPC by optimizing only ~200 parameters projected randomly back into the full weight space. One reading of that result: the pretrained model already has the right features, and fine-tuning just steers them. Hu et al. (2021) took the next step: if the effective update is low-dimensional, why not bake that constraint into the parameterization itself? Don’t sample a random subspace — learn one, by writing the update as a product of two low-rank matrices.
So the hypothesis is: ΔW (the change to a weight matrix during
fine-tuning) is well-approximated by BA, where B ∈ ℝ^{d×r} and
A ∈ ℝ^{r×k} and r is much smaller than d or k.
The empirical result: yes, often it is. r=8 frequently matches r=64 on downstream tasks. The QLoRA paper, in an Alpaca-style sweep on LLaMA-7B with LoRA applied to all layers, reported that rank had effectively no effect on final task performance — a narrower observation than “rank doesn’t matter,” but a striking one.
What this doesn’t prove: it doesn’t show that a full fine-tune “actually” only moves weights in a low-rank subspace. It only shows that if you constrain the update to be low-rank, you don’t lose much. Which is what you wanted operationally; the deeper claim is still under active debate.
The mechanics
For a frozen pretrained weight matrix W₀ ∈ ℝ^{d×k}, the LoRA
parameterization replaces
y = W₀ x
with
y = W₀ x + (α/r) · B A x
where B ∈ ℝ^{d×r}, A ∈ ℝ^{r×k}, and α is a scalar scaling factor.
In the original paper’s scheme, B is initialized to zero and A to
random Gaussian, so BA = 0 at step 0 and the adapted model starts as
an exact copy of the base. (Library defaults differ in details — e.g.
Hugging Face PEFT uses Kaiming-uniform for A and zeros for B — but
the “starts at zero update” property is preserved.) Training updates
only A and B; W₀ never moves.
A few details worth knowing:
- Why two matrices? Rank-
rupdates can be written as the sum ofrouter products.BAis just that, factored. Theris the bottleneck — it’s the only place rank can be lost. - The
α/rscaling factor. It decouples the scale of the update from the choice of rank. The paper itself just callsαa constant they didn’t tune; in practice many recipes tieαtor(oftenα = rorα = 2r) so you can sweep rank without re-tuning the learning rate. The reference implementation (microsoft/LoRA) defaultsα = 1— the “tie alpha to rank” convention is a community norm, not a paper prescription. - Where you put adapters matters. The original paper applied LoRA
only to the attention projection matrices (
W_q,W_v). Later practice often extends adapters to the feed-forward blocks too; QLoRA, for instance, applies LoRA to all linear layers in a transformer block. No clean meta-analysis settles whether “all layers always wins,” so treat this as common practice rather than a settled result.
Why the savings are this dramatic
Stack a few effects:
- Trainable parameters drop ~100–1000×. ~65k vs. ~16.8M per matrix in the example above. Across a whole model, well under 1% of weights are trainable.
- Optimizer state drops by the same factor. Adam’s two extra FP32 numbers per trainable parameter are now negligible. This is the real memory win — bigger than the parameter-count win for most setups, because optimizer state is what blows up VRAM in full fine-tuning.
- Adapters are tiny on disk. A 7B LoRA adapter is on the order of tens of MB; the base model is ~14 GB at FP16. Distributing per-customer fine-tunes becomes feasible.
- No inference penalty if you merge. At serving time you can compute
W = W₀ + (α/r) BAonce and use the merged weights. Same forward-pass cost as the base model. Or you can keep the adapter separate and hot-swap, paying a small overhead per layer.
QLoRA: stacking the trick on quantization
QLoRA (Dettmers et al., May 2023) is the most consequential follow-up. The base model is quantized to 4-bit (NF4, a data type tailored for normally-distributed weights), kept frozen, and dequantized on the fly to 16-bit (FP16/BF16) to compute forward and backward passes. Only the LoRA adapters and the optimizer state for them live in 16-bit. The headline: fine-tune a 65B model on one 48 GB GPU, recovering full 16-bit fine-tuning quality on the benchmarks they tested. A 48 GB card is still a serious purchase, but it’s a single card — which is what moved 65B fine-tuning out of cluster territory. Your five customers now need one machine, not one cluster.
Where it gets subtle
- LoRA is not literally identical to full fine-tuning. A late-2024 paper, LoRA vs Full Fine-tuning: An Illusion of Equivalence, argues the two reach different solutions in weight space even when downstream metrics look similar — different generalization, different forgetting behavior. The practical advice from the LoRA community has long been: treat LoRA as the default but evaluate against full fine-tuning when the task is hard or the data is large. Whether the two are equivalent in any deeper sense is an open dispute rather than a settled question.
- Rank isn’t the only knob. Which layers you adapt, the
αscaling, learning rate, and whether you also tune embeddings all matter. The rank-vs-quality curve is famously flat for many tasks, which is why r=8 keeps working. - It’s parameter-efficient, not knowledge-efficient. LoRA shrinks how many parameters update, not how much data you need. If your fine-tune is bad, more rank won’t save you; better data will.
- Knowledge injection is still hard. As with full fine-tuning, LoRA is a worse tool for stuffing new factual knowledge into a model than for adjusting style, format, or task framing. Retrieval is usually the right answer for “here are the facts I want it to know.”
The whole technique is a bet on a structural observation about fine-tuning, not a clever optimizer trick. That’s why it generalizes — the same idea works for transformer LLMs, diffusion image models, and just about anything else where you start from a strong pretrained base. Your five customers now cost one base model plus five files small enough to sit side by side on a single server.
You started with LoRA = freeze base weights + a low-rank update BA.
What did this post add that the line hides? — + the savings come from the optimizer state, not the parameter count. Adam keeps two extra
FP32 numbers per trainable parameter, so cutting trainable parameters
by ~1000× cuts the thing that was actually blowing up your VRAM. The
tiny adapter files are the pleasant side effect; the memory is the
reason it works on hardware you own.
Check yourself
Before you go — someone reports that bumping LoRA rank from 8 to 128 barely changed their eval scores, and concludes “rank doesn’t matter, always use 8.” What’s wrong with that inference?
Answer
They’ve generalized from one task. The flat rank-vs-quality curve is a real and widely reported observation — QLoRA’s Alpaca-style sweep on LLaMA-7B found rank had effectively no effect on final task performance — but it’s an observation about the tasks people have tested, mostly style-and-format adaptation on modest datasets. The intrinsic-rank hypothesis says the update for that task lives in a low-dimensional subspace; it doesn’t say every conceivable adaptation does. If their fine-tune is small and stylistic, r=8 is a fine default. If it’s large and genuinely teaches new behavior, the honest move is to sweep rank and also compare against a full fine-tune — which is exactly what the Illusion of Equivalence line of work argues for.
Next — a sixth customer arrives. They want the model to answer in formal Japanese, refuse anything outside their product line, and always end with a specific disclaimer. They have 4,000 example conversations. Is this a LoRA job, and what would you predict about rank?
Answer
Yes — this is close to the ideal case. Everything they’re asking for is style, format and task framing: tone, refusal boundary, output template. That’s the kind of adaptation the intrinsic-rank hypothesis covers best, and it’s the kind that has produced the famously flat rank-vs-quality curve. Start at r=8, expect it to be enough, and spend your effort on the 4,000 examples rather than on the rank sweep — LoRA is parameter-efficient, not data-efficient, so the data is the thing that decides whether it works. Ship the adapter alongside the other five; the base model doesn’t change.
And one more — a colleague wants to use LoRA to teach a model your company’s internal product catalog, which changes weekly. Good plan?
Answer
Probably not, for two separate reasons. First, LoRA is parameter-efficient, not knowledge-efficient — it shrinks how many weights update, not how much data you need, and fine-tuning in general is a poor tool for injecting specific facts compared to retrieval. Second, “changes weekly” means you’d be retraining weekly and the model would carry stale facts between runs with no way to invalidate them. Retrieval handles both: put the catalog in a store, fetch the relevant rows at query time, and updating it is a database write. Save the LoRA for teaching the model how to talk about the catalog.
Famous related terms
- Full fine-tuning —
full fine-tuning = unfreeze every weight + run gradient descent on the whole model. The baseline LoRA is compared against. Best quality when you can afford it; rarely worth it. - PEFT —
PEFT = umbrella for "fine-tune by training a tiny number of extra parameters". LoRA is the most popular member; adapters, prefix tuning, IA³, and prompt tuning are siblings. - QLoRA —
QLoRA = 4-bit quantized frozen base + 16-bit LoRA adapters on top. Dettmers et al., 2023. The reason single-GPU fine-tuning of 65B models is on the table. - Adapters (Houlsby et al., 2019) —
adapter = small bottleneck MLP inserted between transformer layers + only train it. The pre-LoRA PEFT idea; LoRA’s main practical advantage is no extra inference latency when you merge. - DoRA —
DoRA ≈ LoRA + decompose into magnitude and direction. A 2024 variant (arXiv:2402.09353) that reports beating LoRA at the same parameter budget on the tasks its authors tested; whether that holds on yours, and whether it’s worth the added complexity, is a “try it” question. - Why fine-tuning is cheap — the broader story LoRA is one ingredient of.
- Why VRAM is the bottleneck — explains why the optimizer-state savings are the win that matters.
- Intrinsic dimension —
intrinsic dimension = the smallest subspace you can fine-tune in and still do well. The Aghajanyan/Li line of work LoRA built its hypothesis on.
Going deeper
- LoRA: Low-Rank Adaptation of Large Language Models (Hu, Shen, Wallis, Allen-Zhu, Li, Wang, Wang, Chen — arXiv 2106.09685, June 2021; ICLR 2022). The primary source; read §4 for exactly which matrices they adapted and what happened when they varied rank.
- The Hugging Face PEFT library — the explainer-by-code: five minutes in the LoRA example answers “what does this actually look like in a training script, and how small is the trainable footprint really?”
- LoRA vs Full Fine-tuning: An Illusion of Equivalence (arXiv 2410.21228, 2024) — the rabbit hole, for “when should I not trust LoRA?” It argues the two land in different places in weight space even when the benchmark numbers match.