Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why is fine-tuning so cheap compared to pretraining?

Pretraining a frontier model costs tens of millions of dollars. Fine-tuning the same model on your data can cost less than a pizza. Why the four-orders-of-magnitude gap?

AI & ML intermediate Apr 29, 2026 · updated Aug 25, 2026 · 15 min read

On this page

The picture version

Six pictures for a reader who has never trained a model. The prose below fills in the seams the pictures skip.

1 · The problem

Two bills for the same model, four orders of magnitude apart.

building the model trillions of words thousands of chips weeks of wall-clock time a CFO asks questions the order of magnitude is public; the digits mostly aren’t teaching it your support voice 2,000 support tickets one machine a coffee break nobody asks same model, same hardware, same optimiser The obvious explanation is that one is just a smaller version of the other. That is the wrong model, and it’s the one this post breaks.
The same 2,000 tickets are the running example throughout. The gap isn’t scale — the two jobs are solving different problems, and nearly all the cost lives in the problem the cheap one never has to solve.

2 · Naive attempt 1

Train on the tickets from scratch and you don’t get a bad bot. You get gibberish.

random numbers 0.41   -0.02   0.88 -0.7   0.13   -0.55 0.09   0.62   -0.31 it knows nothing at all your 2,000 tickets a few hundred thousand words what comes out the the of and the refund of of the a not a bad bot — not a sentence It has never learned that letters make words, that words have roles, that “Paris” and “France” go together, that arguments have shape. every one of those has to be discovered from raw text, and 2,000 tickets is nowhere near enough
Two thousand examples cannot teach a model language. The expensive thing pretraining buys is not knowledge of your product — it is the machinery that makes any text make sense at all, and your tickets never pay for it.

3 · What you are actually renting

The tickets don’t teach new features. They pick which existing ones to combine.

letters clump into words words have roles and grammar facts about the world hang together arguments and code have structure “sound like our support team” trillions of words bought all of this 2,000 tickets bought this the same stack, seen as a bill And you start from a place where the model already handles language sensibly — so a few hundred careful steps finish the job instead of hundreds of thousands.
Pretraining lays down structure that any downstream task can lean on; adaptation only selects and reweights it. You inherit both the features and a sane starting point, which is the second discount and the easiest one to miss.

4 · Naive attempt 2

Update every weight and the job dies — not on the weights, on the optimiser’s notes.

update everything the weights the gradients optimiser note, one per weight optimiser note, one per weight two more tensors the size of the model, often stored more precisely than the weights this is the wall you actually hit only allow a small update the weights — frozen, untouched the update you allow its notes its notes a rank-8 update to a 4096×4096 grid is about 65,000 numbers instead of 16.8 million the wall recedes the passes still run through the whole model, so this saves memory far more than arithmetic
Adam-style optimisers keep two running notes per parameter they update, so a full update needs several copies of the model resident at once. Constrain the update to a low-rank shape and those notes shrink with it — the saving is memory, not FLOPs.

5 · Naive attempt 3

It’s so cheap that people use it on the one job it’s bad at.

good at this how the answer sounds what shape it comes in what the job even is every one of your 2,000 tickets demonstrates all three bad at this this quarter’s pricing last week’s changelog which service is down right now it can memorise them, inefficiently, at a price Push hard on a narrow pile of text and the model can quietly get worse at general things you never meant to touch. so keep the changelog in a document store you can edit on a Tuesday, and look it up instead
Voice, format and task framing are what the tickets actually demonstrate; facts are not. Training facts in risks eroding capabilities you never meant to touch — a risk to manage rather than a certainty, and one that retrieval sidesteps entirely.

6 · Keep this card

No single step is mysterious. They multiply.

fine-tuning = pretraining minus building the model it rents a representation instead of making one and the bill drops on four separate axes at once less data thousands, not trillions fewer steps hundreds, not 100,000s fewer weights moved often under 1% of them smaller notes none for the frozen weights Four large factors, on four different axes, multiplied — which is how you land four orders of magnitude apart.
Picture to keep: pretraining builds the whole factory — machines, wiring, trained staff; your fine-tune walks into the finished factory and changes what’s printed on the labels. Where the picture breaks: changing the labels can quietly damage the machines, which nothing in a real factory does.

Why it exists

Your company’s support bot keeps opening every reply with “Certainly!”, calls your product “the platform” instead of its actual name, and occasionally invents a refund policy you’ve never offered. You have two thousand real support tickets in a database, each paired with the reply a human actually sent. You upload them, start a fine-tuning job, go get coffee, and come back to a model that sounds like your support team. The bill is small enough that nobody asks you to justify it.

That support-ticket fine-tune is the running example for the rest of this post — same two thousand tickets, same model, all the way down.

Now put it next to the other number. Pretraining the frontier LLM you just fine-tuned involved trillions of tokens, thousands of GPUs, weeks of wall-clock time, and a budget large enough to make a CFO ask follow-up questions. Public estimates put the compute bill for the largest models in the tens to low hundreds of millions of dollars — though the exact numbers for any specific frontier model are usually not public, so treat the order of magnitude as the load-bearing fact, not the digits.

Four-plus orders of magnitude separate those two bills. You probably assume the explanation is scale: that fine-tuning is the same operation as pretraining, just on a smaller pile of data. That’s the model to break first. Pretraining and fine-tuning are not one operation at two sizes — they are solving two different problems, and almost all of the cost lives in the problem fine-tuning never has to solve.

Why it matters now

Every team building on top of LLMs makes a fine-tune-vs-prompt-vs-RAG decision regularly, and most of them make it badly because they don’t have a working model of what fine-tuning is for.

If you don’t have a feel for why fine-tuning is cheap, you’ll either under-use it (sticking to prompting when fine-tuning would obviously win) or over-use it (fine-tuning when prompting or retrieval would have been fine).

The short answer

fine-tuning ≈ pretraining minus the part where you build the representations from scratch

Picture to keep: pretraining builds the whole factory — machines, wiring, trained staff; your support-ticket fine-tune walks into the finished factory and changes what’s printed on the labels. The analogy breaks in one place worth remembering: changing the labels can quietly damage the machines. Nothing in a factory works that way, and it’s the subtlety the last third of this post is about.

Pretraining has to teach the model everything: grammar, world facts, reasoning patterns, the geometry of language itself. Fine-tuning gets to assume all of that is already in the weights, and only nudges them toward a narrower behavior. Less data, fewer steps, often only a small fraction of the parameters touched. The expensive thing already happened.

How it works

The cleanest way to see where the money goes is to try to build the support-ticket bot the naive way and watch it fail, three times in a row. Each fix is one of the reasons fine-tuning is cheap.

Naive attempt 1: train a model on your two thousand tickets

Start from random weights, feed in the tickets, run gradient descent.

This doesn’t produce a bad support bot. It produces something that can’t form a sentence. A randomly-initialized network knows nothing. It doesn’t know that letters group into words, that words have parts of speech, that “Paris” and “France” are related, that code has syntax, that arguments have structure. Every one of those facts has to be discovered from scratch, by gradient descent, from raw token streams.

That’s what trillions of tokens of pretraining buys: an internal representation of language and the world good enough that next-token prediction gets sharp. The model ends up with — and this is the load-bearing claim — a set of features in its hidden layers that already encode most of what any downstream task needs. A 2019 paper from Tenney et al. probed BERT’s layers and found that information for classical NLP tasks (parts of speech, parsing, coreference) is laid down across layers in roughly the same order a hand-written pipeline would run them — a softer claim than “BERT is a pipeline,” but enough to make the point: pretraining quietly assembles structure that downstream tasks can lean on. (Later work has pushed back on the strongest version of the pipeline reading; treat it as a suggestive picture, not a settled mechanistic claim.)

The fix: start from someone else’s checkpoint. Once those representations exist, your two thousand tickets aren’t teaching the model language anymore. They’re teaching it which existing features to combine. That’s a much smaller learning problem — and it comes with a second, easily-missed discount. Pretraining starts from random weights, and the standard practitioner’s account of why that’s expensive — high loss everywhere, gradients that don’t point anywhere useful yet, a long warmup before any structure appears — is intuition backed by the shape of training curves, not a proved description of the landscape. Your fine-tune starts somewhere the model already handles language sensibly, so gradient descent behaves: small learning rate, a few hundred to a few thousand steps, done. Weeks of optimizer time become hours — for the same model, on the same hardware; don’t read that ratio across different setups. This is exactly the transfer learning move from image models a decade ago — pretrain on ImageNet, fine-tune on your bird photos — at larger scale.

Naive attempt 2: fine-tune, but update every weight in the model

Now the tickets actually teach the model something. But start the job on one GPU and it dies before the first step, and not because of the weights themselves. The optimizer is the problem: Adam-style optimizers keep running moment estimates for every parameter they update — two tensors the size of the model, commonly kept in higher precision than the weights themselves. Full fine-tuning a large model means holding the weights, the gradients, and those moments resident at once. That’s the wall most people actually hit — memory, not FLOPs.

The fix: only allow a small update in the first place. Here’s the surprising empirical observation that powers modern parameter-efficient fine-tuning: you can often express a useful fine-tune as a very low-rank update on top of frozen base weights, and lose little quality compared to a full fine-tune.

The LoRA paper (Hu et al., arXiv 2021; ICLR 2022) made this concrete. It hypothesized that the effective update needed during adaptation has low “intrinsic rank,” then showed empirically that constraining the update to be a low-rank matrix — say, rank 8 or 16, in a model where the original weight matrix is thousands by thousands — gets you fine-tuning quality close to the full-update baseline on a range of tasks. That’s an enormous compression. A rank-8 update to a 4096×4096 matrix has 4096×8 + 8×4096 ≈ 65k parameters instead of ~16.8M — about 256× fewer.

The interpretation is something like: the pretrained model already lives near the right answer for downstream tasks, and a successful adaptation can usually be written as a small rotation in a few directions rather than a rebuild. (Note what this doesn’t prove: it doesn’t show that a full fine-tune would only have moved a tiny subspace of the weights; it shows that you can get most of the benefit by only allowing a low-rank update in the first place. Whether the intrinsic-rank hypothesis is the right explanation, and how universally it holds across tasks and architectures, is still an active research area.)

What’s solid in practice: LoRA-style updates work well across a wide range of supervised and preference fine-tuning, and a fine-tune that trains under 1% of the parameters can match a full fine-tune for many real tasks. And attempt 2’s wall recedes: you never allocate optimizer moments for the frozen weights, only for the tiny update. That’s usually the real saving, more than the FLOPs. It doesn’t make memory a non-issue — the base weights and the activations are still resident — but it removes the term that was several times the size of the model.

Naive attempt 3: since it’s this cheap, fine-tune in the whole product manual

The bot now sounds right. So you throw the product docs, the pricing page, and last quarter’s changelog into the training set too, reasoning that facts are just more text.

This is the failure mode that can cost teams months. Fine-tuning is a good tool for changing style, format, and task framing — the things your two thousand tickets demonstrate on every line. It’s a poor tool for installing new facts. The network can memorize specific facts during fine-tuning, but it’s inefficient, and pushing hard on a narrow distribution risks eroding the base model’s general capabilities — the catastrophic forgetting problem. The failure you’re courting: a bot that recites last quarter’s pricing page confidently and has gotten measurably worse at the general reasoning you never meant to touch. How badly this bites depends on the learning rate, the data mix, and how long you train — it’s a risk to manage, not a guarantee.

The fix: don’t train the facts in — look them up. For “let the model use this knowledge,” retrieval is the right tool, not fine-tuning. Keep the fine-tune for voice and format; keep the changelog in a document store where you can edit it on a Tuesday without retraining anything. (LoRA is often reached for here too — the base weights are never modified, so the original checkpoint is always recoverable by dropping the adapter. That’s a weaker guarantee than “no forgetting”: the adapted model can still behave worse on things the adapter interferes with.)

Putting the savings together

Stack the discounts and the asymmetry stops being mysterious:

You don’t get a single 10,000× improvement from any one of these. You get a large factor from each, on a different axis, and they multiply — which is how you land four orders of magnitude apart without any one step being mysterious. (The per-axis factors are back-of-envelope, not measured.)

Where it gets subtle

You started with fine-tuning ≈ pretraining minus building the representations from scratch. What did the three failures add? — + you also inherit a good starting point in the loss landscape, and you only need to pay for the parameters you actually let move. That’s why the opening assumption was wrong: pretraining and fine-tuning aren’t one operation at two scales, they’re different operations. Pretraining builds a representation; your support-ticket fine-tune rents one.

Check yourself

Before you go — a colleague proposes fine-tuning the model nightly on that day’s new support tickets so it “stays current” on outages and known bugs. Cheap, right? What goes wrong?

Answer

Two things, and neither is the GPU bill. First, this is attempt 3 in disguise: the thing they want the model to learn is facts (which service is down, which bug is open), and fine-tuning is a bad fact-installer — inefficient, and prone to degrading unrelated capabilities. Retrieval over the ticket database gets today’s outage into the answer without touching a weight. Second, each nightly run pulls the model further onto a narrow distribution; over weeks that compounds into catastrophic forgetting. The cheapness is real, which is exactly why it’s tempting to use the tool on a problem it doesn’t solve.

And one on the money: if you switch your fine-tune from full-parameter to LoRA, which cost drops most — the data you collect, the FLOPs per step, or the GPU memory you need to rent?

Answer

Memory. The dataset is unchanged, and the forward and backward passes still run through the whole model, so the FLOPs per step don’t collapse. What disappears is optimizer state: no Adam moments for the frozen weights, only for a rank-8-ish update. That’s why LoRA is what brings a fine-tune of a small or mid-sized model within reach of a single GPU — it moved the wall you were actually hitting. (Whether it fits your GPU still depends on the base model’s size and your batch and sequence lengths, which LoRA doesn’t touch.)

Going deeper