Why does mixture-of-experts exist?
A 671B-parameter model whose per-token compute is closer to a 37B one. The trick isn't compression — it's that most of the weights sit out most of the time.
On this page
The picture version
Six pictures for a reader who has never thought about what happens inside a chatbot. The prose below fills in the seams the pictures skip.
1 · The problem
You asked for one French word. Every part of the model woke up for it.
2 · The idea
Replace the one big block with many, and put a receptionist in front.
3 · What that buys
Two numbers that used to be one.
4 · The first thing that goes wrong
Left alone, the receptionist sends everyone to the same two desks.
5 · The bill you didn’t see
Only two blocks run — but all of them have to be somewhere reachable.
6 · Keep this card
The whole thing on one index card.
Why it exists
Ask a chatbot “what’s the French word for bread?” and time it. Then ask it to explain a Rust borrow-checker error. The second question is harder in every way a human would measure — and on a dense model, both answers cost roughly the same per word. That French question is the running example for this post. Every weight in the network wakes up for it: the parts that know Rust, the parts that memorized obscure chemistry, the parts that handle Mandarin. All of them, for pain.
That’s the asymmetry at the heart of LLM scaling. Bigger models are smarter — that part has held up depressingly well. But every parameter you add gets paid for twice: once at training time (more flops to push gradients through) and again at inference time (more bytes to stream through GPU memory for every single token). At some point you can’t afford the model you’d actually like to have.
And you suspect, watching it answer pain, that you didn’t need most of it. A dense model — one where every weight participates in every forward pass — is paying full price for capacity it isn’t using on this particular input.
Mixture-of-experts is the architecture that takes that suspicion seriously. Instead of one big FFN block in some or all transformer layers, you have N parallel FFN blocks (“experts”), and a tiny router that, per token, picks the k it thinks are most relevant. The other N − k don’t run. You get the capacity of a much bigger model and the per-token compute of a much smaller one — at the cost of a lot more memory and a lot more engineering.
Why it matters now
The practical stake is per-token price and how much context you can afford. When a provider serves your French question from a sparse model, the flops it bills for are the active parameters, not the total — which is why a lab can ship a model with frontier-scale knowledge at a price that tracks something much smaller. If you’re comparing model costs, capacities, or serving requirements, “active” vs “total” is the number you have to read separately:
- Mixtral 8x7B (Mistral, Jan 2024) — ~47B total parameters, ~13B active per token, two of eight experts routed per layer. (paper)
- DeepSeek-V3 (Dec 2024) — 671B total parameters, 37B active per token, with one shared expert plus 8 of 256 routed experts active. (technical report)
A 671B MoE and a 671B dense model are not the same kind of object — same rough footprint on disk, wildly different active compute, bandwidth profiles, and serving topologies. That difference is why you can’t infer “this model needs a cluster” or “this model will be slow” from a parameter count alone anymore.
(An aside on the most-asked example: GPT-4’s architecture has never been published. There’s a persistent industry rumor, traceable to a 2023 George Hotz remark and repeated widely since, that it’s an MoE with around eight experts. I’m mentioning it only because you’ll meet the claim — it is uncitable speculation, not a fact, and nothing in this post depends on it.)
The short answer
MoE = N expert FFNs + a router that picks k of them per token
Picture to keep: a call center with 256 specialists at desks. Every caller talks to a receptionist who routes them to two desks; the other 254 people sit idle for that call. You still have to pay rent on all 256 desks — that’s the catch — but you only pay for two conversations.
In MoE layers, the feed-forward block is replaced by a bank of N parallel feed-forward sub-networks. A small router (usually a single linear layer plus a top-k selection) reads the token’s hidden state and decides which k experts get to process it. Only those k run. The output is a weighted combination of their results, then the rest of the transformer continues as normal. Training jointly learns the experts and the router. (Not every transformer block has to be an MoE layer — many architectures interleave dense and MoE layers — but for simplicity assume every FFN is replaced.)
How it works
Start with the dense version and break it one constraint at a time.
Naive version. Picture the standard transformer block:
hidden -> attention -> add+norm -> FFN -> add+norm -> hidden'
The FFN is a fat two-layer MLP, and in a dense model it’s where the bulk of the parameters live.
Why it breaks. Your French question streams all of it, every token, every layer. Scale the model up and that cost scales with it, whether or not the extra capacity is relevant to pain.
The fix. MoE swaps that single FFN for N of them in parallel:
hidden -> attention -> add+norm -> router picks k of N FFNs -> combine -> add+norm -> hidden'
Your French question’s tokens now touch a handful of experts per layer instead of all of them. The router for a token x computes scores router(x) -> R^N, takes the top-k (often k=2, sometimes k=1 in Switch Transformer, k=8 in DeepSeek-V3), normalizes those k scores into weights (the exact normalization varies — softmax is common; DeepSeek-V3 uses sigmoid affinities normalized over the selected experts), runs only those k experts, and combines their outputs by the weighted sum. The other N − k experts contribute nothing for this token. Different tokens take different paths through the same layer. That’s the whole idea.
A few things follow from that picture, and most of MoE engineering is dealing with them.
Why “active” vs “total” parameters
The router activates a fixed k of N experts per token. So per-token compute scales with k, not N. The capacity of the model — how much knowledge it can encode — scales with N. So you decouple the two:
- Total params ≈ what fits in GPU memory and what the model can know.
- Active params ≈ what runs per token and what you pay for in flops.
DeepSeek-V3 is the cleanest demonstration: 671B total / 37B active is roughly an 18× ratio. You’re getting the knowledge-soaking capacity of something near 671B with the per-token compute closer to 37B. Whether the quality matches a hypothetical dense 671B is a different question, and there’s no public like-for-like dense-671B baseline to compare against. The honest claim is “much more capacity per active flop than dense,” not “free lunch.”
Break #1: the router collapses
The first thing that goes wrong if you’re not careful: the router collapses. A few experts win every routing decision, the rest are never picked, never get gradient, and atrophy. You’ve trained an N-expert model that effectively uses 2.
The fix is an auxiliary load-balancing loss added to training: a penalty term that nudges the router toward using all experts roughly equally over a batch. Different MoE papers use different versions (Switch Transformer’s auxiliary load-balancing loss; DeepSeek-V3’s auxiliary-loss-free balancing scheme plus a small sequence-wise balance loss). They don’t fully solve the problem — a lot of recent MoE routing research is in schemes that balance better with less explicit pressure — but they keep training from degenerating into “one expert does everything.”
There’s also capacity factor: each expert can only process so many tokens per batch. If too many tokens want the same expert in the same step, somebody loses. Many systems set a capacity factor and drop overflowed tokens; others engineer the routing and balancing carefully enough to avoid drops in practice.
Break #2: the hidden bill is memory
Here’s the seam in the “MoE is cheaper” story. Active params are what scales the compute per token. But for inference you have to keep all N experts somewhere in the serving system — on a single GPU, across a multi-GPU node, or sharded across hosts — because the router might pick any of them on any token. So while MoE reduces active compute compared to an equally large dense model, it adds parameter-residency and cross-device routing costs that the memory bandwidth wall of LLM serving already makes painful.
This is why MoE inference is mostly a story about distributed serving:
- Expert parallelism — different experts live on different GPUs. Tokens get routed across the network to wherever their experts are. This works at scale but introduces all-to-all communication on every MoE layer, which is its own performance cliff.
- Big-host inference — a single machine with enough VRAM (or a tightly-coupled multi-GPU node) holds the whole model. Cleaner, but the hardware is expensive.
- Offloading — keep cold experts on CPU memory or NVMe, page them in on demand. My read of why this stays niche: with a large batch, the tokens collectively want most of the experts anyway, so there’s nothing cold left to offload, and a miss costs you a storage round-trip in the middle of a forward pass. There’s no public benchmark pinning down where the crossover sits.
So the MoE pitch isn’t “cheaper to deploy.” It’s a conditional claim: potentially cheaper per quality unit, if you can afford the extra memory and the routing overhead. Whether that math works out is a question about your traffic pattern and your hardware, not a property of the architecture.
Break #3: the intuitions the name gives you are wrong
A few things that surprise people the first time:
- Experts don’t specialize the way the name suggests. This is where the call-center picture breaks, and it’s worth knowing: there is probably no “French desk” your question gets routed to. The Mixtral paper includes a routing analysis that looks for topic-based expert specialization and reports no obvious pattern of it; the structure they do surface looks positional and syntactic rather than subject-matter-based. I’d treat “experts fire on domains” as the wrong default picture — though some MoE work (DeepSeekMoE, for example) deliberately pushes toward finer-grained specialization, so it isn’t a law. No survey settles how much category-like structure emerges across architectures.
- MoE training instability is real. Sparse routing creates discrete, non-differentiable decisions in the forward pass (you literally drop experts), and the gradients flowing through the router are noisy. Sparse-gated MoE was introduced by Shazeer et al. in 2017; Switch Transformer (2021) simplified routing to top-1 and reported stable training at much larger scale. A lot of post-2021 MoE work is about taming this further.
- Batching can help less. In a dense model, batching multiple requests amortizes weight-loading costs across them. In an MoE, different tokens in the batch route to different experts, so the same expert might be needed for only a fraction of the batch — fragmenting the work. The “free lunch” of batch-amortized memory bandwidth shrinks. This is part of why MoE inference engines look different from dense ones.
You started with MoE = N expert FFNs + a router that picks k of them per token. What did this post add that the line hides? — + the savings are in flops, not bytes. All N experts still have to be resident somewhere in the serving system, because the router might call any of them on the next token. So MoE moves the bill rather than shrinking it: less compute per token, more memory and more network. Your French question is cheap to answer; the machine that can answer it is not cheap to own.
Check yourself
Before you go — someone points at DeepSeek-V3’s “671B total / 37B active” and says the model runs on the same hardware as a 37B dense model. What have they got wrong?
Answer
They’ve confused active compute with resident memory. 37B active means the flops per token are in 37B territory — that’s the real win. But the router can send the next token to any of the 256 experts in a layer, so all 671B parameters have to be somewhere the serving system can reach in time: one very large host, a multi-GPU node, or sharded across machines with all-to-all communication on every MoE layer. A 37B dense model fits on hardware this does not. The right summary is “37B of work on a 671B footprint.”
And one more — you’re serving an MoE and you increase batch size, expecting the usual throughput win from amortizing weight loads. Would you predict the same speedup as on a dense model?
Answer
No, expect less. On a dense model, every request in the batch needs every weight, so loading a weight once and using it 64 times is a straight 64× amortization. On an MoE, the tokens in your batch scatter across experts — a given expert might be needed by a small fraction of the batch, so you pay to load its weights and then do proportionally little work with them. The amortization is fragmented. (There’s a second-order effect worth noticing: as the batch grows, the batch collectively wants more of the experts, which is why offloading cold experts stops working at exactly the batch sizes where you wanted throughput.)
Famous related terms
- Sparsely-gated mixture-of-experts —
sparse MoE = dense MoE + top-k routing— the Shazeer et al. 2017 paper that introduced the modern formulation: thousands of experts, only a few active per example, applied between LSTM layers. Pre-transformer, but it laid down the sparse-gating template widely reused in later work. - Switch Transformer —
Switch ≈ MoE with k=1— Fedus, Zoph, Shazeer (2021). Routes each token to exactly one expert, simpler and more stable than top-2. The paper reports training sparse models up to a trillion parameters. - Mixtral 8x7B —
Mixtral = 8 experts + top-2 routing in MoE layers— Mistral’s first MoE release (paper). 47B total, 13B active. Note: not “8 separate 7B models bolted together” — only the FFN blocks are replicated; attention layers are shared. - DeepSeekMoE / DeepSeek-V3 —
DeepSeekMoE = many fine-grained experts + shared experts + balancing tweaks— pushes toward a much larger N (256 routed + 1 shared) with a small k, aiming to increase combinatorial flexibility and isolate shared/common knowledge in dedicated experts. (V3 report) - Router / gating network —
router = small linear layer + top-k selection per token. The (usually) tiny layer that picks experts per token. The single most fragile component in an MoE; much of the routing research is about making it better-behaved. - Active vs total parameters —
active params = what runs per token;total params = what fits in memory. The distinction the whole architecture exists to create. When you see “37B active / 671B total,” now you know what it’s saying. - Expert parallelism —
EP = "shard the experts across GPUs"— the distributed-systems half of MoE serving, and where the all-to-all communication shows up. - Auxiliary load-balancing loss —
aux loss = main loss + a penalty that nudges the router toward using all experts. The training-time hack that keeps the router from collapsing onto a few favorites. Every MoE paper has its own version.
Going deeper
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer — Shazeer et al., 2017. The primary source: read it for why routing needs a balancing loss at all, which is the failure the whole design orbits. (Pre-transformer — the experts sit between LSTM layers — but the skeleton is the one still in use.)
- Mixtral of Experts — Jiang et al., 2024. The explainer: a complete, readable MoE design in one place, plus their attempt to find topic specialization in the router and what they found instead.
- DeepSeek-V3 Technical Report — DeepSeek-AI, 2024. The rabbit hole, for “what does it actually take to train and serve one of these?” Unusually candid about the engineering.
On GPT-4: there is no official architecture paper, and the widely repeated “8-expert MoE” figure traces to a 2023 George Hotz remark that OpenAI has never confirmed. There’s nothing to link that would settle it.