Why is fine-tuning so cheap compared to pretraining?
Pretraining a frontier model costs tens of millions of dollars. Fine-tuning the same model on your data can cost less than a pizza. Why the four-orders-of-magnitude gap?
On this page
- The picture version
- Why it exists
- Why it matters now
- The short answer
- How it works
- Naive attempt 1: train a model on your two thousand tickets
- Naive attempt 2: fine-tune, but update every weight in the model
- Naive attempt 3: since it’s this cheap, fine-tune in the whole product manual
- Putting the savings together
- Where it gets subtle
- Check yourself
- Famous related terms
- Going deeper
The picture version
Six pictures for a reader who has never trained a model. The prose below fills in the seams the pictures skip.
1 · The problem
Two bills for the same model, four orders of magnitude apart.
2 · Naive attempt 1
Train on the tickets from scratch and you don’t get a bad bot. You get gibberish.
3 · What you are actually renting
The tickets don’t teach new features. They pick which existing ones to combine.
4 · Naive attempt 2
Update every weight and the job dies — not on the weights, on the optimiser’s notes.
5 · Naive attempt 3
It’s so cheap that people use it on the one job it’s bad at.
6 · Keep this card
No single step is mysterious. They multiply.
Why it exists
Your company’s support bot keeps opening every reply with “Certainly!”, calls your product “the platform” instead of its actual name, and occasionally invents a refund policy you’ve never offered. You have two thousand real support tickets in a database, each paired with the reply a human actually sent. You upload them, start a fine-tuning job, go get coffee, and come back to a model that sounds like your support team. The bill is small enough that nobody asks you to justify it.
That support-ticket fine-tune is the running example for the rest of this post — same two thousand tickets, same model, all the way down.
Now put it next to the other number. Pretraining the frontier LLM you just fine-tuned involved trillions of tokens, thousands of GPUs, weeks of wall-clock time, and a budget large enough to make a CFO ask follow-up questions. Public estimates put the compute bill for the largest models in the tens to low hundreds of millions of dollars — though the exact numbers for any specific frontier model are usually not public, so treat the order of magnitude as the load-bearing fact, not the digits.
Four-plus orders of magnitude separate those two bills. You probably assume the explanation is scale: that fine-tuning is the same operation as pretraining, just on a smaller pile of data. That’s the model to break first. Pretraining and fine-tuning are not one operation at two sizes — they are solving two different problems, and almost all of the cost lives in the problem fine-tuning never has to solve.
Why it matters now
Every team building on top of LLMs makes a fine-tune-vs-prompt-vs-RAG decision regularly, and most of them make it badly because they don’t have a working model of what fine-tuning is for.
- Buy-vs-build math. “Should we fine-tune our own model?” sounds capital-intensive. In compute terms it usually isn’t — for a job like the support-ticket fine-tune, the budget line that tends to hurt is the weeks spent cleaning and labelling those tickets and building an eval, not the GPU hours.
- Why LoRA and friends took over. Parameter-efficient fine-tuning methods turn an already-cheap operation into an even cheaper one. They exist because someone noticed that the update you’re trying to learn during fine-tuning has a very particular shape, and you don’t need to touch every weight to express it.
- Open-weight ecosystems. Hugging Face is packed with thousands of fine-tunes of a handful of base models. For small and mid-sized open models, a single GPU and a weekend is genuinely enough to produce one — that’s only possible because the cost curve drops off a cliff after pretraining. (It stops being true as you scale up the base model.)
- The “post-training” stack. Modern instruction-tuned and aligned models (most chat models you actually use) are produced by running several fine-tuning stages on top of a pretrained checkpoint — supervised fine-tuning, then preference optimization like RLHF or DPO. At frontier labs these stages typically do keep updating the full set of base weights rather than freezing them, though open-model practice often uses parameter-efficient methods instead. The whole pipeline only makes economic sense because each stage after pretraining is dramatically cheaper than the stage before.
If you don’t have a feel for why fine-tuning is cheap, you’ll either under-use it (sticking to prompting when fine-tuning would obviously win) or over-use it (fine-tuning when prompting or retrieval would have been fine).
The short answer
fine-tuning ≈ pretraining minus the part where you build the representations from scratch
Picture to keep: pretraining builds the whole factory — machines, wiring, trained staff; your support-ticket fine-tune walks into the finished factory and changes what’s printed on the labels. The analogy breaks in one place worth remembering: changing the labels can quietly damage the machines. Nothing in a factory works that way, and it’s the subtlety the last third of this post is about.
Pretraining has to teach the model everything: grammar, world facts, reasoning patterns, the geometry of language itself. Fine-tuning gets to assume all of that is already in the weights, and only nudges them toward a narrower behavior. Less data, fewer steps, often only a small fraction of the parameters touched. The expensive thing already happened.
How it works
The cleanest way to see where the money goes is to try to build the support-ticket bot the naive way and watch it fail, three times in a row. Each fix is one of the reasons fine-tuning is cheap.
Naive attempt 1: train a model on your two thousand tickets
Start from random weights, feed in the tickets, run gradient descent.
This doesn’t produce a bad support bot. It produces something that can’t form a sentence. A randomly-initialized network knows nothing. It doesn’t know that letters group into words, that words have parts of speech, that “Paris” and “France” are related, that code has syntax, that arguments have structure. Every one of those facts has to be discovered from scratch, by gradient descent, from raw token streams.
That’s what trillions of tokens of pretraining buys: an internal representation of language and the world good enough that next-token prediction gets sharp. The model ends up with — and this is the load-bearing claim — a set of features in its hidden layers that already encode most of what any downstream task needs. A 2019 paper from Tenney et al. probed BERT’s layers and found that information for classical NLP tasks (parts of speech, parsing, coreference) is laid down across layers in roughly the same order a hand-written pipeline would run them — a softer claim than “BERT is a pipeline,” but enough to make the point: pretraining quietly assembles structure that downstream tasks can lean on. (Later work has pushed back on the strongest version of the pipeline reading; treat it as a suggestive picture, not a settled mechanistic claim.)
The fix: start from someone else’s checkpoint. Once those representations exist, your two thousand tickets aren’t teaching the model language anymore. They’re teaching it which existing features to combine. That’s a much smaller learning problem — and it comes with a second, easily-missed discount. Pretraining starts from random weights, and the standard practitioner’s account of why that’s expensive — high loss everywhere, gradients that don’t point anywhere useful yet, a long warmup before any structure appears — is intuition backed by the shape of training curves, not a proved description of the landscape. Your fine-tune starts somewhere the model already handles language sensibly, so gradient descent behaves: small learning rate, a few hundred to a few thousand steps, done. Weeks of optimizer time become hours — for the same model, on the same hardware; don’t read that ratio across different setups. This is exactly the transfer learning move from image models a decade ago — pretrain on ImageNet, fine-tune on your bird photos — at larger scale.
Naive attempt 2: fine-tune, but update every weight in the model
Now the tickets actually teach the model something. But start the job on one GPU and it dies before the first step, and not because of the weights themselves. The optimizer is the problem: Adam-style optimizers keep running moment estimates for every parameter they update — two tensors the size of the model, commonly kept in higher precision than the weights themselves. Full fine-tuning a large model means holding the weights, the gradients, and those moments resident at once. That’s the wall most people actually hit — memory, not FLOPs.
The fix: only allow a small update in the first place. Here’s the surprising empirical observation that powers modern parameter-efficient fine-tuning: you can often express a useful fine-tune as a very low-rank update on top of frozen base weights, and lose little quality compared to a full fine-tune.
The LoRA paper (Hu et al., arXiv 2021; ICLR 2022) made this concrete. It hypothesized that the effective update needed during adaptation has low “intrinsic rank,” then showed empirically that constraining the update to be a low-rank matrix — say, rank 8 or 16, in a model where the original weight matrix is thousands by thousands — gets you fine-tuning quality close to the full-update baseline on a range of tasks. That’s an enormous compression. A rank-8 update to a 4096×4096 matrix has 4096×8 + 8×4096 ≈ 65k parameters instead of ~16.8M — about 256× fewer.
The interpretation is something like: the pretrained model already lives near the right answer for downstream tasks, and a successful adaptation can usually be written as a small rotation in a few directions rather than a rebuild. (Note what this doesn’t prove: it doesn’t show that a full fine-tune would only have moved a tiny subspace of the weights; it shows that you can get most of the benefit by only allowing a low-rank update in the first place. Whether the intrinsic-rank hypothesis is the right explanation, and how universally it holds across tasks and architectures, is still an active research area.)
What’s solid in practice: LoRA-style updates work well across a wide range of supervised and preference fine-tuning, and a fine-tune that trains under 1% of the parameters can match a full fine-tune for many real tasks. And attempt 2’s wall recedes: you never allocate optimizer moments for the frozen weights, only for the tiny update. That’s usually the real saving, more than the FLOPs. It doesn’t make memory a non-issue — the base weights and the activations are still resident — but it removes the term that was several times the size of the model.
Naive attempt 3: since it’s this cheap, fine-tune in the whole product manual
The bot now sounds right. So you throw the product docs, the pricing page, and last quarter’s changelog into the training set too, reasoning that facts are just more text.
This is the failure mode that can cost teams months. Fine-tuning is a good tool for changing style, format, and task framing — the things your two thousand tickets demonstrate on every line. It’s a poor tool for installing new facts. The network can memorize specific facts during fine-tuning, but it’s inefficient, and pushing hard on a narrow distribution risks eroding the base model’s general capabilities — the catastrophic forgetting problem. The failure you’re courting: a bot that recites last quarter’s pricing page confidently and has gotten measurably worse at the general reasoning you never meant to touch. How badly this bites depends on the learning rate, the data mix, and how long you train — it’s a risk to manage, not a guarantee.
The fix: don’t train the facts in — look them up. For “let the model use this knowledge,” retrieval is the right tool, not fine-tuning. Keep the fine-tune for voice and format; keep the changelog in a document store where you can edit it on a Tuesday without retraining anything. (LoRA is often reached for here too — the base weights are never modified, so the original checkpoint is always recoverable by dropping the adapter. That’s a weaker guarantee than “no forgetting”: the adapted model can still behave worse on things the adapter interferes with.)
Putting the savings together
Stack the discounts and the asymmetry stops being mysterious:
- Less data. Pretraining ingests trillions of tokens. A typical supervised fine-tune sees thousands to a few million examples — even generously normalized to tokens, that’s many orders of magnitude less data through the optimizer. (The exact factor depends heavily on example length and which fine-tune you’re talking about; the load-bearing fact is just “way less.”)
- Fewer steps. A pretraining run is hundreds of thousands of optimizer steps; a fine-tune is hundreds to a few thousand. (Illustrative magnitudes — the actual counts vary a lot by run.)
- Fewer parameters touched. Full fine-tune updates 100% of weights; LoRA-style methods often update under 1%.
- Smaller optimizer state. With LoRA you don’t have to store full-precision Adam moments for the frozen weights, which is often the real memory bottleneck in full fine-tuning.
You don’t get a single 10,000× improvement from any one of these. You get a large factor from each, on a different axis, and they multiply — which is how you land four orders of magnitude apart without any one step being mysterious. (The per-axis factors are back-of-envelope, not measured.)
Where it gets subtle
- “Cheap to run” is not “cheap to do well.” GPU-hour cost is the thing that drops by orders of magnitude. The expensive part of a serious fine-tuning project is now data quality and evaluation. In my experience the fine-tunes that disappoint in production fail on the dataset, not on compute — but no public study quantifies that split, so treat it as a practitioner’s rule of thumb rather than a measured finding.
- The frontier-lab numbers really are mostly opaque. I’m comfortable with the claim that there’s a multi-order-of-magnitude gap between pretraining and fine-tuning compute for a given model. I’m not comfortable putting a precise dollar number on either side without citing a specific public estimate. Take any specific figure you read — including the “tens of millions” framing in the opener — as back-of-envelope.
You started with fine-tuning ≈ pretraining minus building the representations from scratch. What did the three failures add? — + you also inherit a good starting point in the loss landscape, and you only need to pay for the parameters you actually let move. That’s why the
opening assumption was wrong: pretraining and fine-tuning aren’t one
operation at two scales, they’re different operations. Pretraining builds a
representation; your support-ticket fine-tune rents one.
Check yourself
Before you go — a colleague proposes fine-tuning the model nightly on that day’s new support tickets so it “stays current” on outages and known bugs. Cheap, right? What goes wrong?
Answer
Two things, and neither is the GPU bill. First, this is attempt 3 in disguise: the thing they want the model to learn is facts (which service is down, which bug is open), and fine-tuning is a bad fact-installer — inefficient, and prone to degrading unrelated capabilities. Retrieval over the ticket database gets today’s outage into the answer without touching a weight. Second, each nightly run pulls the model further onto a narrow distribution; over weeks that compounds into catastrophic forgetting. The cheapness is real, which is exactly why it’s tempting to use the tool on a problem it doesn’t solve.
And one on the money: if you switch your fine-tune from full-parameter to LoRA, which cost drops most — the data you collect, the FLOPs per step, or the GPU memory you need to rent?
Answer
Memory. The dataset is unchanged, and the forward and backward passes still run through the whole model, so the FLOPs per step don’t collapse. What disappears is optimizer state: no Adam moments for the frozen weights, only for a rank-8-ish update. That’s why LoRA is what brings a fine-tune of a small or mid-sized model within reach of a single GPU — it moved the wall you were actually hitting. (Whether it fits your GPU still depends on the base model’s size and your batch and sequence lengths, which LoRA doesn’t touch.)
Famous related terms
- Pretraining —
pretraining = neural net + "predict the next token" objective + internet-scale corpus. The expensive stage that builds the base representations. - Supervised fine-tuning (SFT) —
SFT = pretrained model + (input, desired output) pairs + a few epochs of gradient descent. The simplest fine-tuning recipe; the baseline against which everything else is compared. - LoRA —
LoRA = freeze base weights + add a low-rank update + only train that update. The technique that made parameter-efficient fine-tuning the default, and the reason attempt 2’s memory wall stopped mattering. - PEFT —
PEFT ≈ umbrella term for "fine-tune by training a tiny number of extra parameters instead of touching the base weights". Includes LoRA, adapters, prefix tuning, IA³, and others. - RLHF / DPO —
RLHF = SFT + reward model + RL loop;DPO = SFT + direct preference optimization on pairs. Preference-based post-training stages applied after supervised fine-tuning. Both inherit the cheapness for the same reasons described here. - In-context learning —
in-context learning = prompt + frozen weights + a continuation that happens to be the task answer. The cheaper alternative to fine-tuning when you only have a handful of examples and don’t want to update weights at all. - Transfer learning —
transfer learning = pretrain on a big general task + fine-tune on your specific small one. The general principle that LLM fine-tuning is one instance of. - LLM —
LLM = neural net + "predict the next token" objective at scale. The thing being fine-tuned.
Going deeper
- LoRA: Low-Rank Adaptation of Large Language Models (Hu et al., arXiv 2021; ICLR 2022) — the primary source, if you want to see exactly what the intrinsic-rank hypothesis claims and what the experiments do and don’t establish.
- The Hugging Face PEFT library docs — the fastest answer to “what does a cheap fine-tune actually look like as code,” and how small the trainable footprint really is.
- Rabbit hole: How transferable are features in deep neural networks? (Yosinski et al., 2014) — answers “which layers actually transfer, and what does freezing them cost you,” a decade before anyone said LoRA.