Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why does mixture-of-experts exist?

A 671B-parameter model whose per-token compute is closer to a 37B one. The trick isn't compression — it's that most of the weights sit out most of the time.

AI & ML intermediate Apr 29, 2026 · updated Aug 25, 2026 · 14 min read

On this page

The picture version

Six pictures for a reader who has never thought about what happens inside a chatbot. The prose below fills in the seams the pictures skip.

1 · The problem

You asked for one French word. Every part of the model woke up for it.

“what’s the French word for bread?” the part thatknows Rust the part thatknows chemistry the part thatknows Mandarin the part thatknows French awakeawakeawakeawake All of it, for pain. and a hard question about a borrow-checker error costs the machine about the same per word every extra parameter is paid for twice: once while training, and again on every token you ever generate
In an ordinary model every weight participates in every answer. You are paying full price for capacity this particular question doesn’t need — and that suspicion is the one this architecture takes seriously.

2 · The idea

Replace the one big block with many, and put a receptionist in front.

before one big block everything runs after the router reads the token, scores every block these two run for this token the rest sit out Different tokens take different paths through the same layer. That is the whole idea.
The one fat block inside each layer becomes many parallel blocks, and a tiny learned router picks a couple of them per token. The ones it doesn’t pick simply don’t run — and the next token may take a completely different route.

3 · What that buys

Two numbers that used to be one.

one published model, two very different numbers 671B total — what it can know, and what must fit in memory 37B active — what runs per token, and what you pay for The capacity grows with the number of blocks. The bill grows with how many you switch on. so “how big is this model?” stopped being one question — you have to read both numbers
Because a fixed few blocks run per token, per-token compute stops tracking the model’s size. Capacity and cost, which used to be the same dial, come apart — though whether the quality matches a same-size dense model is a separate question with no public like-for-like comparison.

4 · The first thing that goes wrong

Left alone, the receptionist sends everyone to the same two desks.

what happens if you don’t intervene how often each block gets picked the winners get all the practice, the rest never improve and wither with a balancing penalty in training a nudge that costs the model something whenever it plays favourites now all the blocks earn their keep
A router with nothing keeping it honest collapses onto a handful of favourites, and the unpicked blocks never receive the training signal that would make them worth picking. Every design in this family ships some version of a balancing pressure — and none of them solves it completely.

5 · The bill you didn’t see

Only two blocks run — but all of them have to be somewhere reachable.

every block, sitting in expensive memory, all the time the two that ran for this token — the router could pick any of the others for the next one what you saved arithmetic per token the real win, and a large one what you now pay memory, and traffic between chips blocks live on different machines, so tokens travel So it isn’t cheaper to own. It is cheaper to run, if you can afford to own it.
The router might call any block on the next token, so every one of them has to stay somewhere the system can reach in time. The saving is in arithmetic, not in bytes — which moves the bill from compute to memory and network rather than removing it.

6 · Keep this card

The whole thing on one index card.

mixture of experts = many parallel blocks + a router that picks a few, per token + and the saving is in arithmetic, not bytes your French question is cheap to answer; the machine that can answer it is not cheap to own
Picture to keep: a call centre with 256 specialists at desks, where every caller is routed to two of them and the other 254 sit idle — you still pay rent on all 256. Where the picture breaks, and it matters: there is probably no “French desk.” When researchers went looking for topic-based specialisation they found structure that looks positional and syntactic instead, so don’t picture experts as subject-matter departments.

Why it exists

Ask a chatbot “what’s the French word for bread?” and time it. Then ask it to explain a Rust borrow-checker error. The second question is harder in every way a human would measure — and on a dense model, both answers cost roughly the same per word. That French question is the running example for this post. Every weight in the network wakes up for it: the parts that know Rust, the parts that memorized obscure chemistry, the parts that handle Mandarin. All of them, for pain.

That’s the asymmetry at the heart of LLM scaling. Bigger models are smarter — that part has held up depressingly well. But every parameter you add gets paid for twice: once at training time (more flops to push gradients through) and again at inference time (more bytes to stream through GPU memory for every single token). At some point you can’t afford the model you’d actually like to have.

And you suspect, watching it answer pain, that you didn’t need most of it. A dense model — one where every weight participates in every forward pass — is paying full price for capacity it isn’t using on this particular input.

Mixture-of-experts is the architecture that takes that suspicion seriously. Instead of one big FFN block in some or all transformer layers, you have N parallel FFN blocks (“experts”), and a tiny router that, per token, picks the k it thinks are most relevant. The other N − k don’t run. You get the capacity of a much bigger model and the per-token compute of a much smaller one — at the cost of a lot more memory and a lot more engineering.

Why it matters now

The practical stake is per-token price and how much context you can afford. When a provider serves your French question from a sparse model, the flops it bills for are the active parameters, not the total — which is why a lab can ship a model with frontier-scale knowledge at a price that tracks something much smaller. If you’re comparing model costs, capacities, or serving requirements, “active” vs “total” is the number you have to read separately:

A 671B MoE and a 671B dense model are not the same kind of object — same rough footprint on disk, wildly different active compute, bandwidth profiles, and serving topologies. That difference is why you can’t infer “this model needs a cluster” or “this model will be slow” from a parameter count alone anymore.

(An aside on the most-asked example: GPT-4’s architecture has never been published. There’s a persistent industry rumor, traceable to a 2023 George Hotz remark and repeated widely since, that it’s an MoE with around eight experts. I’m mentioning it only because you’ll meet the claim — it is uncitable speculation, not a fact, and nothing in this post depends on it.)

The short answer

MoE = N expert FFNs + a router that picks k of them per token

Picture to keep: a call center with 256 specialists at desks. Every caller talks to a receptionist who routes them to two desks; the other 254 people sit idle for that call. You still have to pay rent on all 256 desks — that’s the catch — but you only pay for two conversations.

In MoE layers, the feed-forward block is replaced by a bank of N parallel feed-forward sub-networks. A small router (usually a single linear layer plus a top-k selection) reads the token’s hidden state and decides which k experts get to process it. Only those k run. The output is a weighted combination of their results, then the rest of the transformer continues as normal. Training jointly learns the experts and the router. (Not every transformer block has to be an MoE layer — many architectures interleave dense and MoE layers — but for simplicity assume every FFN is replaced.)

How it works

Start with the dense version and break it one constraint at a time.

Naive version. Picture the standard transformer block:

hidden -> attention -> add+norm -> FFN -> add+norm -> hidden'

The FFN is a fat two-layer MLP, and in a dense model it’s where the bulk of the parameters live.

Why it breaks. Your French question streams all of it, every token, every layer. Scale the model up and that cost scales with it, whether or not the extra capacity is relevant to pain.

The fix. MoE swaps that single FFN for N of them in parallel:

hidden -> attention -> add+norm -> router picks k of N FFNs -> combine -> add+norm -> hidden'

Your French question’s tokens now touch a handful of experts per layer instead of all of them. The router for a token x computes scores router(x) -> R^N, takes the top-k (often k=2, sometimes k=1 in Switch Transformer, k=8 in DeepSeek-V3), normalizes those k scores into weights (the exact normalization varies — softmax is common; DeepSeek-V3 uses sigmoid affinities normalized over the selected experts), runs only those k experts, and combines their outputs by the weighted sum. The other N − k experts contribute nothing for this token. Different tokens take different paths through the same layer. That’s the whole idea.

A few things follow from that picture, and most of MoE engineering is dealing with them.

Why “active” vs “total” parameters

The router activates a fixed k of N experts per token. So per-token compute scales with k, not N. The capacity of the model — how much knowledge it can encode — scales with N. So you decouple the two:

DeepSeek-V3 is the cleanest demonstration: 671B total / 37B active is roughly an 18× ratio. You’re getting the knowledge-soaking capacity of something near 671B with the per-token compute closer to 37B. Whether the quality matches a hypothetical dense 671B is a different question, and there’s no public like-for-like dense-671B baseline to compare against. The honest claim is “much more capacity per active flop than dense,” not “free lunch.”

Break #1: the router collapses

The first thing that goes wrong if you’re not careful: the router collapses. A few experts win every routing decision, the rest are never picked, never get gradient, and atrophy. You’ve trained an N-expert model that effectively uses 2.

The fix is an auxiliary load-balancing loss added to training: a penalty term that nudges the router toward using all experts roughly equally over a batch. Different MoE papers use different versions (Switch Transformer’s auxiliary load-balancing loss; DeepSeek-V3’s auxiliary-loss-free balancing scheme plus a small sequence-wise balance loss). They don’t fully solve the problem — a lot of recent MoE routing research is in schemes that balance better with less explicit pressure — but they keep training from degenerating into “one expert does everything.”

There’s also capacity factor: each expert can only process so many tokens per batch. If too many tokens want the same expert in the same step, somebody loses. Many systems set a capacity factor and drop overflowed tokens; others engineer the routing and balancing carefully enough to avoid drops in practice.

Break #2: the hidden bill is memory

Here’s the seam in the “MoE is cheaper” story. Active params are what scales the compute per token. But for inference you have to keep all N experts somewhere in the serving system — on a single GPU, across a multi-GPU node, or sharded across hosts — because the router might pick any of them on any token. So while MoE reduces active compute compared to an equally large dense model, it adds parameter-residency and cross-device routing costs that the memory bandwidth wall of LLM serving already makes painful.

This is why MoE inference is mostly a story about distributed serving:

So the MoE pitch isn’t “cheaper to deploy.” It’s a conditional claim: potentially cheaper per quality unit, if you can afford the extra memory and the routing overhead. Whether that math works out is a question about your traffic pattern and your hardware, not a property of the architecture.

Break #3: the intuitions the name gives you are wrong

A few things that surprise people the first time:

You started with MoE = N expert FFNs + a router that picks k of them per token. What did this post add that the line hides? — + the savings are in flops, not bytes. All N experts still have to be resident somewhere in the serving system, because the router might call any of them on the next token. So MoE moves the bill rather than shrinking it: less compute per token, more memory and more network. Your French question is cheap to answer; the machine that can answer it is not cheap to own.

Check yourself

Before you go — someone points at DeepSeek-V3’s “671B total / 37B active” and says the model runs on the same hardware as a 37B dense model. What have they got wrong?

Answer

They’ve confused active compute with resident memory. 37B active means the flops per token are in 37B territory — that’s the real win. But the router can send the next token to any of the 256 experts in a layer, so all 671B parameters have to be somewhere the serving system can reach in time: one very large host, a multi-GPU node, or sharded across machines with all-to-all communication on every MoE layer. A 37B dense model fits on hardware this does not. The right summary is “37B of work on a 671B footprint.”

And one more — you’re serving an MoE and you increase batch size, expecting the usual throughput win from amortizing weight loads. Would you predict the same speedup as on a dense model?

Answer

No, expect less. On a dense model, every request in the batch needs every weight, so loading a weight once and using it 64 times is a straight 64× amortization. On an MoE, the tokens in your batch scatter across experts — a given expert might be needed by a small fraction of the batch, so you pay to load its weights and then do proportionally little work with them. The amortization is fragmented. (There’s a second-order effect worth noticing: as the batch grows, the batch collectively wants more of the experts, which is why offloading cold experts stops working at exactly the batch sizes where you wanted throughput.)

Going deeper

On GPT-4: there is no official architecture paper, and the widely repeated “8-expert MoE” figure traces to a 2023 George Hotz remark that OpenAI has never confirmed. There’s nothing to link that would settle it.