Why model merging works at all
Take two fine-tunes of the same model, average their weights element-wise, and you often get a model better than either parent. Naively, this shouldn't work — neural net loss surfaces are wildly non-convex. The reason it works tells you something deep about where fine-tuning actually lives.
On this page
The picture version
Six pictures for a reader who has never trained or merged anything. The prose below fills in the seams the pictures skip.
1 · The absurd move
Add two model files together, divide by two, upload the result.
2 · Why it shouldn’t work
Halfway between two valleys is usually a ridge.
3 · The false premise
These aren’t two separate answers. They’re two tents in one valley.
4 · The load-bearing clause
Same parent, and the line between them stays low. Different parents, and it doesn’t.
5 · Where it stops working
Three ways to leave the valley.
6 · Keep this card
Four questions, in order, before you merge anything.
Why it exists
Everyone has met the moment where two files that do almost the same thing are sitting next to each other and you think: could I just… combine them? Usually the answer is no. Here is a case where the answer is yes, and it’s stranger than it sounds. Scroll far enough down an open-model leaderboard and you start hitting names with suffixes like -slerp, -ties, -dare, -merge-v3. Nobody trained those. Somebody took two model files that already existed, combined their weights with a script that runs in minutes, and uploaded the result — and it sits above models that cost real money to train. The first time you notice this it looks like a leaderboard exploit. It isn’t; it’s a real technique, and the reason it works is worth understanding.
Here’s the setup, and the running example for the rest of this post. You have two fine-tuned versions of the same base LLM — one tuned for code, one tuned for medical Q&A. Each is a giant tensor of weights. You want a model that’s good at both. The textbook moves are: train a third model on a mixture of both datasets, or use a router that picks one expert per query, or fine-tune on top of one of them with the other’s data.
Now consider the absurd move: just take the two weight tensors and average them, element by element. merged = (model_A + model_B) / 2. No retraining, no router, no extra data. Just arithmetic on the parameters of two completely separate fine-tunes.
The naive prediction is that this should produce garbage. Neural networks are highly non-linear functions of their weights. The loss surface they’re trained on is famously non-convex — that is, full of hills, valleys, and saddle points, where the midpoint between two low-loss configurations could easily be a high-loss configuration. Average two reasonable solutions to a hard optimization problem and, in the general case, you get an unreasonable one. That’s why nobody used to bother trying.
But empirically, on fine-tunes that share a pretrained base, weight averaging works. Wortsman et al.’s “Model soups” paper (2022, arXiv:2203.05482) showed that averaging the weights of dozens of independently fine-tuned CLIP models produced a single model that beat every individual ingredient on ImageNet, with no extra inference cost. Since then, the trick has spread: task-vector arithmetic, TIES-merging, DARE, and a sprawling ecosystem of merge recipes on Hugging Face. How often merges outrank their own ingredients on public leaderboards isn’t something anyone tracks, and leaderboard composition changes fast, so treat that part as unmeasured — but Wortsman et al.’s CLIP result is the measured version of the same claim.
The interesting question is why this works. The answer turns out to be specific and load-bearing: fine-tunes don’t actually go very far from where they started.
Why it matters now
The concrete present-day use is search. Merging is cheap enough that it changes what you try: if you already have ten fine-tunes, you can generate a hundred candidate merges, evaluate them, and ship the best — at the cost of running evals, not of training runs. That’s the shift, and it’s what your code-plus-medical model is going to come out of. What share of top open-weight models are merges is not something anyone publishes, so there’s no number to put on it.
Merging is also adjacent to two other areas, with a caveat worth stating. Continual learning wants to add a capability without retraining, and merging is a direct tool for that. Federated learning also averages weights, but its classic result (FedAvg) is iterative averaging during a coordinated training run — the models are re-synchronized constantly, which is a stronger condition than “two fine-tunes that share a parent.” Same arithmetic, different guarantee; don’t read results from one as evidence for the other.
The short answer
model merging = element-wise weight averaging of fine-tunes that share a pretrained starting point
Picture to keep: a valley that pretraining spent millions of dollars finding. Your code model and your medical model are two tents pitched a short walk apart on the valley floor — so the midpoint between them is still valley floor, not a cliff. Like that, except the valley has thousands of dimensions and you can’t see it; the only evidence you’re still inside it is that the loss doesn’t spike along the path.
It works because fine-tuning, despite the name, doesn’t actually move the model very far. Fine-tunes from the same pretrained checkpoint stay inside a roughly flat, connected region of the loss surface — the same loss basin their parent lived in. The straight line between two points inside one basin stays inside the basin. So the average isn’t a leap into the void; it’s a step inside the neighborhood the fine-tunes were already exploring.
How it works
Follow the naive prediction and watch where it turns out to be wrong.
Naive prediction: averaging two solutions to a non-convex problem should give you a bad solution. What actually happens: for your code model and your medical model, it doesn’t. So one of the premises must be false — and the false one is “these are two independent solutions to a hard optimization problem.” They aren’t. Here’s why, in three steps.
1. Fine-tuning is a small perturbation
Pretraining a frontier LLM involves trillions of tokens and weeks on thousands of GPUs. Producing your code model or your medical model involves, typically, a few thousand to a few million tokens for a few hours. The gradients are smaller, the learning rate is lower, and you stop early. Whatever the metaphor “fine-tuning” suggests, the standard account is that the weights move very little relative to their own scale — small enough that the two fine-tunes end up as near neighbours.
This is the same territory LoRA lives in: constraining the fine-tuning update to a low-rank subspace costs surprisingly little on the tasks people have tested. Be careful with the direction of that inference, though — LoRA shows the update can be small and low-rank without much loss, not that an unconstrained full fine-tune necessarily is. The two claims are related and often conflated; only the first is well-supported.
2. Linear mode connectivity
Frankle, Dziugaite, Roy, and Carbin (2020, arXiv:1912.05671) — building on earlier work by Garipov, Izmailov, Podoprikhin, Vetrov, and Wilson on loss-surface geometry — showed empirically, in the settings they studied, what’s now called linear mode connectivity: two networks trained from the same initialization (with the same data ordering up to some point, then diverging) tend to be connected by a straight line of low loss in weight space. It’s a robust empirical finding, not a theorem. You can interpolate between them and the loss along the path doesn’t spike.
This is the load-bearing fact for merging. If A and B are linearly mode-connected, then (A + B)/2 has loss comparable to A and B — not somewhere on the other side of a barrier. The mean is in the basin.
The shared-initialization condition turns out to be the crucial caveat. Two networks trained from different random inits typically don’t connect linearly; the line between them passes through high-loss regions. This is why you can’t just merge any two models — the recipe explicitly requires a shared parent.
3. Pretraining as the basin selector
Here’s the synthesis. Pretraining is enormously expensive partly because it does the work of finding a deep, wide loss basin in the absurdly high-dimensional weight space — a basin that generalizes well, where many directions of small perturbation still produce a working language model. Fine-tuning then does the much smaller job of relocating to a particular point inside that basin that happens to be good at the fine-tuning task.
When you have two fine-tunes sharing a pretrained parent, you have two points inside the same basin. Linear mode connectivity says the straight line between them stays in the basin. Averaging is just picking the midpoint of that line. The averaged model inherits whatever properties the basin has — including, often, something close to the union of what the two endpoints learned. The usual story for why is that the directions encoding “knows about code” and “knows about medical Q&A” are largely orthogonal at this scale, so the two updates don’t fight. Flag that as interpretation, not established mechanism: the mode-connectivity literature explains why the midpoint isn’t broken, but it doesn’t establish the orthogonality claim, and it hasn’t been measured convincingly at LLM scale.
What goes wrong, and where the seams are
The story above is clean. Reality is messier:
- Permutation symmetry. Neural networks have a vast symmetry group — you can permute the neurons in a hidden layer (and correspondingly permute the rows of the next layer’s weight matrix) without changing the function. Two networks trained from different inits may compute similar functions but live in different “permutation copies” of the same basin. Naive averaging then destroys both. There’s a research line (Ainsworth, Hayase, Srinivasa 2022, Git Re-Basin) on permutation-aligning networks before averaging, with partial success on small-scale models.
- Interference between fine-tunes. When two fine-tunes both modify the same parameters in opposite directions, averaging cancels them. TIES-merging (Yadav et al., 2023) and DARE (Yu et al., 2023) are recipes that explicitly handle this — drop small or conflicting updates, keep the dominant ones. Both papers report beating plain averaging on the multi-task merges they evaluate; I’d treat that as strong evidence for the mechanism (sign conflicts are real and worth resolving) without assuming either recipe wins on every merge.
- Catastrophic forgetting at the boundary. A fine-tune that drifts far enough from the pretrained parent may have left the original basin, and merging those doesn’t get the linear-mode-connectivity property. Heavily RL-post-trained reasoning models are the obvious suspect — much more optimization pressure than an SFT run. To be clear, “SFT merges cleanly, aggressive RL doesn’t” is community folklore, with no study behind it yet. Treat it as a hypothesis to test on your own checkpoints.
- Capability superposition is fragile. The “orthogonal directions add up” intuition is approximate. In practice, merging too many fine-tunes degrades performance on each — the basin is connected but it isn’t infinite. Most successful merge recipes top out at a small number of ingredients, or use weighted combinations rather than uniform averages.
- It’s not magic. A merged model is bounded above by what’s reachable inside the basin. It can’t acquire a capability neither parent had. (You can sometimes appear to — but that’s usually because the capability was latent in the pretrained base and a parent unlocked it; the merge inherits the unlock.)
So the honest version of “model merging works” is: weight-space averaging of shared-parent fine-tunes is a real, useful, surprisingly cheap operation that exploits a specific geometric fact about the loss surface around pretrained checkpoints. It’s not arbitrage. The pretraining run did the expensive work — finding the basin — and merging just rearranges what’s inside it. That’s also the answer to the leaderboard from the hook: the merge didn’t get something for nothing, it got something for cheap, by spending someone else’s pretraining run.
You started with model merging = element-wise weight averaging of fine-tunes that share a pretrained starting point. What did this post add that the line hides? — + "share a pretrained starting point" is the entire load-bearing clause. Drop it and the same arithmetic produces garbage, because two models from different random inits sit in different basins — or in different permutation copies of the same one. Merging isn’t a property of averaging; it’s a property of a shared parent.
If you want one portable test to carry out of here, it’s four questions in order: same parent? (if no, stop — this isn’t the tool). How far did each one drift? (light SFT is the safe case; heavy RL post-training is the risky one). How many ingredients? (few; quality degrades as you pile them on). Do the updates fight? (if they do, reach for a conflict-aware recipe like TIES rather than a uniform average). Everything else in this post is the reasoning behind those four.
What’s still open
The “fine-tunes share a basin” framing is well-supported empirically but the theoretical understanding is partial. The exact width of the basin, why pretraining produces wide basins (versus narrow ones that wouldn’t tolerate this), and the precise scaling of merge quality with number of ingredients — these are active research questions, not settled.
Nor does the field have a verdict on which merge recipe is best. Linear averaging, TIES, DARE, task arithmetic, SLERP, and various weighted hybrids each have papers showing they win on some benchmark, and nobody has run the like-for-like comparison that would settle it. So the defensible claim is that a conflict-aware recipe often beats naive averaging on multi-task merges — not that any one recipe dominates.
Check yourself
Before you go — you want a model that’s good at code and medical Q&A, and you have two candidates for the pair. Candidate A: two fine-tunes of the same Llama-3-8B base checkpoint, from different teams. Candidate B: a Llama-3-8B code fine-tune and a Mistral-7B medical fine-tune, both excellent. Which pair would you merge, and why?
Answer
Candidate A, and it isn’t close. The shared pretrained parent is the whole mechanism: linear mode connectivity holds for networks that started from the same initialization, so the straight line between two Llama-3-8B fine-tunes stays at low loss. Candidate B’s two models never shared an initialization — different architectures, different pretraining runs — so the line between them passes through high-loss territory, and they’re not even the same shape. (Even with matching shapes and different random inits, permutation symmetry alone can wreck naive averaging: two networks can compute similar functions while their hidden neurons are in different orders.) For Candidate B, the tools are routing or distillation, not arithmetic.
And one more — a team merges eight fine-tunes and reports the result is worse on every individual task than any single ingredient. Is this evidence that merging doesn’t work, or is it the expected outcome?
Answer
Expected, and it’s the seam worth internalizing. “Orthogonal updates compose by addition” is an approximation that degrades as you add ingredients — the basin is connected but finite, and with eight fine-tunes you get more parameters where updates conflict outright and more dilution of each one. This is exactly the failure that TIES-merging and DARE were built for: trim the small updates, resolve the sign conflicts, and keep the dominant ones rather than averaging everything uniformly. The team’s options are fewer ingredients, weighted rather than uniform averaging, or a conflict-aware recipe. It’s not evidence against merging; it’s evidence they exceeded what uniform averaging can carry.
Famous related terms
- Linear mode connectivity —
LMC = property that two networks are connected by a low-loss straight line in weight space— the geometric fact that licenses merging. - Task arithmetic —
task vector = (fine-tuned weights) − (pretrained weights); compose tasks by adding/subtracting these vectors— Ilharco et al. 2022, Editing Models with Task Arithmetic. The “merging is just addition” framing made explicit. - Model soups —
model soup = average of many fine-tunes of the same architecture— Wortsman et al. 2022; the paper that popularized the technique at scale. - TIES-merging —
TIES = trim small updates + resolve sign conflicts + average survivors— Yadav et al. 2023, the standard “do better than uniform averaging” recipe. - LoRA — adjacent story: also exploits “fine-tuning is small,” but compresses the update rather than averaging multiple updates.
- Permutation symmetry —
permutation symmetry = swap neurons in a layer + the function is unchanged— the obstacle to merging across random inits. Solved partially by Git Re-Basin.
Going deeper
- Wortsman et al., Model soups (2022) — the primary source; read it for “does this actually work, measured properly?” Their CLIP/ImageNet results are still the cleanest demonstration.
- Frankle et al., Linear Mode Connectivity and the Lottery Ticket Hypothesis (2020) — answers “why doesn’t the midpoint fall off a cliff?”, which is the geometric fact everything else rests on. Pair with the earlier Garipov et al. Loss Surfaces, Mode Connectivity, and Fast Ensembling of DNNs (2018).
- Ilharco et al., Editing Models with Task Arithmetic (2022) — the rabbit hole: if merging is addition, can you subtract a capability? They try it.
A gap worth naming: there isn’t a good beginner-facing explainer for merging to send someone to. The Model Soups paper is readable enough to stand in for one.