Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why model merging works at all

Take two fine-tunes of the same model, average their weights element-wise, and you often get a model better than either parent. Naively, this shouldn't work — neural net loss surfaces are wildly non-convex. The reason it works tells you something deep about where fine-tuning actually lives.

AI & ML intermediate May 4, 2026 · updated Aug 25, 2026 · 14 min read

On this page

The picture version

Six pictures for a reader who has never trained or merged anything. The prose below fills in the seams the pictures skip.

1 · The absurd move

Add two model files together, divide by two, upload the result.

tuned for code 0.41   -0.02   0.88 -0.7   0.13   -0.55 + tuned for medicine 0.39   0.04   0.81 -0.6   0.11   -0.61 ÷ 2 good at both 0.40   0.01   0.85 -0.65   0.12   -0.58 No retraining. No extra data. Just arithmetic, number by number. it runs in minutes, and the result can land above models that cost real money to train This should not work.
Two fine-tunes of the same base model, averaged element by element, often give you something good at both jobs. Your code model and your medical model are the running example — and the first honest reaction to this recipe is that it ought to produce nonsense.

2 · Why it shouldn’t work

Halfway between two valleys is usually a ridge.

one good model another good model their midpoint Average two answers to a rugged problem and you normally get a bad one.
The surface a network is trained on is full of hills and valleys, so the point halfway between two good solutions can easily be a terrible one. That is the standard reasoning, and it is why nobody used to bother trying — so one of its premises has to be wrong.

3 · The false premise

These aren’t two separate answers. They’re two tents in one valley.

the valley pretraining spent millions finding where pretraining stopped code medicine Fine-tuning, despite the name, barely moves the model. a few hours against weeks; a small learning rate; you stop early So the midpoint isn’t a leap across a ridge. It’s a step between two neighbouring tents.
Both fine-tunes began at the same expensively-found starting point and travelled a short distance from it. They were never two independent answers to a hard problem — which is exactly the premise the pessimistic argument in the last scene depended on.

4 · The load-bearing clause

Same parent, and the line between them stays low. Different parents, and it doesn’t.

same starting point the whole line stays on the floor so the average is a working model different starting points the line climbs straight over a barrier so the average is garbage Merging isn’t a property of averaging. It’s a property of a shared parent.
Networks that started from the same initialisation tend to be joined by a straight path of low loss; networks from different ones are not. It’s a robust empirical finding rather than a theorem, and it is the single fact the whole recipe rests on.

5 · Where it stops working

Three ways to leave the valley.

same function, shuffled A B C C A B you can swap a layer’s neurons around without changing what the model does average them and you destroy both updates that fight same weight one tune pushes it left, the other right averaging cancels them both to nothing too many ingredients the valley is connected, not infinite pile on enough and every task gets worse The recipes that beat plain averaging all attack the middle one: trim the small updates, resolve the conflicts, keep the dominant ones
Merging fails when the two models aren’t really in the same place, when their changes pull the same weight in opposite directions, or when you stack so many that each one is diluted. None of these is a reason the technique doesn’t work — they are the boundary of where the shared-parent argument still holds.

6 · Keep this card

Four questions, in order, before you merge anything.

before you merge, ask: 1  Same parent? If no, stop — wrong tool. 2  How far did each drift? Light is safe. 3  How many ingredients? Few. 4  Do the updates fight? Then don’t just average. The merge didn’t get something for nothing. It got something cheap, by spending someone else’s training run. and it can’t give you a skill neither parent had — the pretraining run did the expensive work of finding the valley
Picture to keep: a valley that pretraining spent millions of dollars finding, with your two models pitched as tents a short walk apart on its floor — so the midpoint is still valley floor, not a cliff. Except the valley has thousands of dimensions and you can’t see it; the only evidence you’re still inside is that nothing gets worse along the path.

Why it exists

Everyone has met the moment where two files that do almost the same thing are sitting next to each other and you think: could I just… combine them? Usually the answer is no. Here is a case where the answer is yes, and it’s stranger than it sounds. Scroll far enough down an open-model leaderboard and you start hitting names with suffixes like -slerp, -ties, -dare, -merge-v3. Nobody trained those. Somebody took two model files that already existed, combined their weights with a script that runs in minutes, and uploaded the result — and it sits above models that cost real money to train. The first time you notice this it looks like a leaderboard exploit. It isn’t; it’s a real technique, and the reason it works is worth understanding.

Here’s the setup, and the running example for the rest of this post. You have two fine-tuned versions of the same base LLM — one tuned for code, one tuned for medical Q&A. Each is a giant tensor of weights. You want a model that’s good at both. The textbook moves are: train a third model on a mixture of both datasets, or use a router that picks one expert per query, or fine-tune on top of one of them with the other’s data.

Now consider the absurd move: just take the two weight tensors and average them, element by element. merged = (model_A + model_B) / 2. No retraining, no router, no extra data. Just arithmetic on the parameters of two completely separate fine-tunes.

The naive prediction is that this should produce garbage. Neural networks are highly non-linear functions of their weights. The loss surface they’re trained on is famously non-convex — that is, full of hills, valleys, and saddle points, where the midpoint between two low-loss configurations could easily be a high-loss configuration. Average two reasonable solutions to a hard optimization problem and, in the general case, you get an unreasonable one. That’s why nobody used to bother trying.

But empirically, on fine-tunes that share a pretrained base, weight averaging works. Wortsman et al.’s “Model soups” paper (2022, arXiv:2203.05482) showed that averaging the weights of dozens of independently fine-tuned CLIP models produced a single model that beat every individual ingredient on ImageNet, with no extra inference cost. Since then, the trick has spread: task-vector arithmetic, TIES-merging, DARE, and a sprawling ecosystem of merge recipes on Hugging Face. How often merges outrank their own ingredients on public leaderboards isn’t something anyone tracks, and leaderboard composition changes fast, so treat that part as unmeasured — but Wortsman et al.’s CLIP result is the measured version of the same claim.

The interesting question is why this works. The answer turns out to be specific and load-bearing: fine-tunes don’t actually go very far from where they started.

Why it matters now

The concrete present-day use is search. Merging is cheap enough that it changes what you try: if you already have ten fine-tunes, you can generate a hundred candidate merges, evaluate them, and ship the best — at the cost of running evals, not of training runs. That’s the shift, and it’s what your code-plus-medical model is going to come out of. What share of top open-weight models are merges is not something anyone publishes, so there’s no number to put on it.

Merging is also adjacent to two other areas, with a caveat worth stating. Continual learning wants to add a capability without retraining, and merging is a direct tool for that. Federated learning also averages weights, but its classic result (FedAvg) is iterative averaging during a coordinated training run — the models are re-synchronized constantly, which is a stronger condition than “two fine-tunes that share a parent.” Same arithmetic, different guarantee; don’t read results from one as evidence for the other.

The short answer

model merging = element-wise weight averaging of fine-tunes that share a pretrained starting point

Picture to keep: a valley that pretraining spent millions of dollars finding. Your code model and your medical model are two tents pitched a short walk apart on the valley floor — so the midpoint between them is still valley floor, not a cliff. Like that, except the valley has thousands of dimensions and you can’t see it; the only evidence you’re still inside it is that the loss doesn’t spike along the path.

It works because fine-tuning, despite the name, doesn’t actually move the model very far. Fine-tunes from the same pretrained checkpoint stay inside a roughly flat, connected region of the loss surface — the same loss basin their parent lived in. The straight line between two points inside one basin stays inside the basin. So the average isn’t a leap into the void; it’s a step inside the neighborhood the fine-tunes were already exploring.

How it works

Follow the naive prediction and watch where it turns out to be wrong.

Naive prediction: averaging two solutions to a non-convex problem should give you a bad solution. What actually happens: for your code model and your medical model, it doesn’t. So one of the premises must be false — and the false one is “these are two independent solutions to a hard optimization problem.” They aren’t. Here’s why, in three steps.

1. Fine-tuning is a small perturbation

Pretraining a frontier LLM involves trillions of tokens and weeks on thousands of GPUs. Producing your code model or your medical model involves, typically, a few thousand to a few million tokens for a few hours. The gradients are smaller, the learning rate is lower, and you stop early. Whatever the metaphor “fine-tuning” suggests, the standard account is that the weights move very little relative to their own scale — small enough that the two fine-tunes end up as near neighbours.

This is the same territory LoRA lives in: constraining the fine-tuning update to a low-rank subspace costs surprisingly little on the tasks people have tested. Be careful with the direction of that inference, though — LoRA shows the update can be small and low-rank without much loss, not that an unconstrained full fine-tune necessarily is. The two claims are related and often conflated; only the first is well-supported.

2. Linear mode connectivity

Frankle, Dziugaite, Roy, and Carbin (2020, arXiv:1912.05671) — building on earlier work by Garipov, Izmailov, Podoprikhin, Vetrov, and Wilson on loss-surface geometry — showed empirically, in the settings they studied, what’s now called linear mode connectivity: two networks trained from the same initialization (with the same data ordering up to some point, then diverging) tend to be connected by a straight line of low loss in weight space. It’s a robust empirical finding, not a theorem. You can interpolate between them and the loss along the path doesn’t spike.

This is the load-bearing fact for merging. If A and B are linearly mode-connected, then (A + B)/2 has loss comparable to A and B — not somewhere on the other side of a barrier. The mean is in the basin.

The shared-initialization condition turns out to be the crucial caveat. Two networks trained from different random inits typically don’t connect linearly; the line between them passes through high-loss regions. This is why you can’t just merge any two models — the recipe explicitly requires a shared parent.

3. Pretraining as the basin selector

Here’s the synthesis. Pretraining is enormously expensive partly because it does the work of finding a deep, wide loss basin in the absurdly high-dimensional weight space — a basin that generalizes well, where many directions of small perturbation still produce a working language model. Fine-tuning then does the much smaller job of relocating to a particular point inside that basin that happens to be good at the fine-tuning task.

When you have two fine-tunes sharing a pretrained parent, you have two points inside the same basin. Linear mode connectivity says the straight line between them stays in the basin. Averaging is just picking the midpoint of that line. The averaged model inherits whatever properties the basin has — including, often, something close to the union of what the two endpoints learned. The usual story for why is that the directions encoding “knows about code” and “knows about medical Q&A” are largely orthogonal at this scale, so the two updates don’t fight. Flag that as interpretation, not established mechanism: the mode-connectivity literature explains why the midpoint isn’t broken, but it doesn’t establish the orthogonality claim, and it hasn’t been measured convincingly at LLM scale.

What goes wrong, and where the seams are

The story above is clean. Reality is messier:

So the honest version of “model merging works” is: weight-space averaging of shared-parent fine-tunes is a real, useful, surprisingly cheap operation that exploits a specific geometric fact about the loss surface around pretrained checkpoints. It’s not arbitrage. The pretraining run did the expensive work — finding the basin — and merging just rearranges what’s inside it. That’s also the answer to the leaderboard from the hook: the merge didn’t get something for nothing, it got something for cheap, by spending someone else’s pretraining run.

You started with model merging = element-wise weight averaging of fine-tunes that share a pretrained starting point. What did this post add that the line hides? — + "share a pretrained starting point" is the entire load-bearing clause. Drop it and the same arithmetic produces garbage, because two models from different random inits sit in different basins — or in different permutation copies of the same one. Merging isn’t a property of averaging; it’s a property of a shared parent.

If you want one portable test to carry out of here, it’s four questions in order: same parent? (if no, stop — this isn’t the tool). How far did each one drift? (light SFT is the safe case; heavy RL post-training is the risky one). How many ingredients? (few; quality degrades as you pile them on). Do the updates fight? (if they do, reach for a conflict-aware recipe like TIES rather than a uniform average). Everything else in this post is the reasoning behind those four.

What’s still open

The “fine-tunes share a basin” framing is well-supported empirically but the theoretical understanding is partial. The exact width of the basin, why pretraining produces wide basins (versus narrow ones that wouldn’t tolerate this), and the precise scaling of merge quality with number of ingredients — these are active research questions, not settled.

Nor does the field have a verdict on which merge recipe is best. Linear averaging, TIES, DARE, task arithmetic, SLERP, and various weighted hybrids each have papers showing they win on some benchmark, and nobody has run the like-for-like comparison that would settle it. So the defensible claim is that a conflict-aware recipe often beats naive averaging on multi-task merges — not that any one recipe dominates.

Check yourself

Before you go — you want a model that’s good at code and medical Q&A, and you have two candidates for the pair. Candidate A: two fine-tunes of the same Llama-3-8B base checkpoint, from different teams. Candidate B: a Llama-3-8B code fine-tune and a Mistral-7B medical fine-tune, both excellent. Which pair would you merge, and why?

Answer

Candidate A, and it isn’t close. The shared pretrained parent is the whole mechanism: linear mode connectivity holds for networks that started from the same initialization, so the straight line between two Llama-3-8B fine-tunes stays at low loss. Candidate B’s two models never shared an initialization — different architectures, different pretraining runs — so the line between them passes through high-loss territory, and they’re not even the same shape. (Even with matching shapes and different random inits, permutation symmetry alone can wreck naive averaging: two networks can compute similar functions while their hidden neurons are in different orders.) For Candidate B, the tools are routing or distillation, not arithmetic.

And one more — a team merges eight fine-tunes and reports the result is worse on every individual task than any single ingredient. Is this evidence that merging doesn’t work, or is it the expected outcome?

Answer

Expected, and it’s the seam worth internalizing. “Orthogonal updates compose by addition” is an approximation that degrades as you add ingredients — the basin is connected but finite, and with eight fine-tunes you get more parameters where updates conflict outright and more dilution of each one. This is exactly the failure that TIES-merging and DARE were built for: trim the small updates, resolve the sign conflicts, and keep the dominant ones rather than averaging everything uniformly. The team’s options are fewer ingredients, weighted rather than uniform averaging, or a conflict-aware recipe. It’s not evidence against merging; it’s evidence they exceeded what uniform averaging can carry.

Going deeper

A gap worth naming: there isn’t a good beginner-facing explainer for merging to send someone to. The Model Soups paper is readable enough to stand in for one.