Why SwiGLU replaced ReLU in transformers
Modern LLMs ditched the simplest activation function in deep learning for a multiplicative gate nobody can fully explain. Here's why.
On this page
The picture version
Five pictures for a reader who has never opened a transformer’s source. The prose below fills in the seams the pictures skip.
1 · The problem
The textbook promised two matrices. The code has three.
2 · Complaint one
Below zero, ReLU stops teaching.
3 · The move
Let one branch decide how much of the other gets through.
4 · The seam
The paper that settled it declined to say why it works.
5 · Keep this card
The whole thing on one index card.
Why it exists
Open the source of any recent open-weights model — Llama, Qwen, Mistral — and
scroll to the
FFN
block inside one transformer layer. That’s our
running example for this post: one FFN block, in one layer, of one modern
model. You expect two weight matrices, because that’s what every textbook
and every diagram of a transformer shows: project up, apply a nonlinearity,
project back down. Instead you find three — gate_proj, up_proj,
down_proj — and an elementwise multiplication sitting between them that no
textbook mentioned.
You’d assume a change like that came from theory: someone proved the old activation was leaving something on the table. It’s closer to the opposite. The three-matrix block won an empirical bake-off, and the author of the paper that established the result explicitly declined to explain why it works. The interesting question isn’t “what is SwiGLU” — it’s how a field ends up standardizing on something nobody can derive.
If you opened a deep-learning textbook in 2016, the activation function on every page
was ReLU.
It is the simplest nonlinearity that works: a kink at zero, a straight line above
it, no exponentials. ResNets used it. The original Transformer paper used it.
For a long time it was the default for the same reason printf is the default
debug tool — it’s cheap, it’s understood, it gets out of the way.
Then something quietly happened between 2017 and 2022. BERT and GPT-2 swapped ReLU for GeLU. Then PaLM and LLaMA swapped GeLU for SwiGLU. By 2024, opening most decoder-only open-weights models, you’d find a feed-forward layer with three matrices instead of two and a multiplication you’d never seen in a textbook. Something pushed a large part of the field to abandon a thing that was famously fine.
Why it matters now
Most modern LLM serving stacks — vLLM and TensorRT-LLM
among them — ship a fused kernel for the gated feed-forward block, because
the three-matrix shape is what the Llama, Qwen, and Mistral families all use.
It isn’t universal: Google’s Gemma line, for instance, gates with GeLU rather
than Swish — same three-matrix shape, different activation on the gate branch.
Quantization schemes and tensor-parallel splits have to handle that shape
specifically. If you’re reading LLaMA or Qwen or Mistral source code and the
mlp block has gate_proj, up_proj, and down_proj, that’s a gated FFN —
in those models specifically, SwiGLU. Knowing why that shape won is the
difference between “the FFN is a black box” and “I can predict how this
kernel allocates memory.”
The short answer
SwiGLU = Swish(xW) ⊙ (xV) → multiplied through W₂
Picture to keep: a mixing desk. One branch carries the signal; the other branch is a row of faders, one per channel, that the network learns to set — and the two get multiplied together before anything leaves the block. ReLU only had a single on/off switch per channel; SwiGLU has a continuous fader whose position depends on the input. (Where the analogy breaks: on a real mixing desk a human sets the faders once; here both branches are computed fresh from the same input on every token, so “signal” and “faders” aren’t fixed roles — the split is a convention of the formula, not a fact about which branch means what.)
In words: instead of the classic feed-forward layer f(xW₁) · W₂, you compute
two projections of the input, pass one through a smooth activation
(Swish),
multiply them elementwise (the “gate”), and then project back down with a
third matrix W₂. The gate lets the network decide, per coordinate, how much
of the other branch’s signal to let through.
How it works
The three matrices in your Llama layer are the residue of a chain of complaints. Each step below is somebody’s fix for the previous step’s annoyance — which is why you can’t derive the shape from first principles, only retrace it.
The classic FFN block (Vaswani et al., 2017). Given input x, compute
h = ReLU(x W₁ + b₁), then y = h W₂ + b₂. Two matmuls, one nonlinearity. The
hidden dimension is typically 4× the model dimension — that’s where most of the
parameter count of a transformer actually lives, more than attention.
What breaks in the classic block: ReLU has two annoyances.
Step 1: ReLU → GeLU (around 2018). Below zero it is
exactly flat — the gradient is zero, so a neuron stuck in the negative region
gets no learning signal (“dying ReLU”). And the kink at zero means the function
isn’t differentiable there, which is fine in practice but ugly in theory.
Hendrycks & Gimpel’s GeLU paper (arXiv:1606.08415) proposed x · Φ(x) where
Φ is the Gaussian CDF.
Same shape as ReLU at the extremes, but smooth, and crucially its gradient
isn’t identically zero for negative inputs, so a neuron sitting just below
zero still gets a learning signal. (It does still decay toward zero far out on
the negative side — the fix is to the flat region near zero, not to the tail.) BERT and GPT-2 picked it up and it became the new default. The
empirical gains were small but consistent.
Step 2: GeLU → SwiGLU (around 2020-2022). This is the weirder jump. Noam Shazeer’s “GLU Variants Improve Transformer” (arXiv:2002.05202, 2020) tested a family of GLU variants in the FFN sublayer. The pattern is:
SwiGLU(x) = (Swish(x W) ⊙ (x V)) W₂
Two input projections (W and V) instead of one. One goes through Swish, the
other stays linear. They get multiplied elementwise (⊙). Then W₂ projects
back. The gate is the new ingredient — the linear branch can amplify or
suppress each coordinate of the activated branch.
Why “free performance”? To keep the parameter count comparable to the
classic two-matrix FFN, SwiGLU implementations shrink the hidden dimension by
2/3 (so roughly 8/3 · d_model instead of 4 · d_model). With matched
parameters, SwiGLU still outperforms GeLU and ReLU on perplexity and
downstream benchmarks. PaLM used it, LLaMA 1/2/3 use it, and most open-weights
models published after late 2022 use it.
Why does the gate help? Honestly — and this is the seam — nobody really knows. Shazeer’s paper closes with a now-famous line attributing the improvement to “divine benevolence.” The hand-wavy story is that multiplicative interactions let the network represent things linear-plus-ReLU can’t easily represent (e.g. quadratic functions of the input), and the gate gives it a learnable per-coordinate dial. But the honest answer is that we have an empirical result that replicates across scales, and a mechanism story that’s plausible but not proven. The field adopted it because the loss curves were better, not because someone derived it from first principles.
One more thing — the cost. SwiGLU has three big matmuls in the FFN
(gate_proj, up_proj, down_proj) instead of two. That changes how kernels
fuse, how TP
splits the weights, and how quantization schemes group the channels. If you’ve
ever wondered why the major inference engines ship a dedicated fused kernel for
this block, that’s why — and it’s why the block in your Llama layer has three names in it
instead of two.
You started with SwiGLU = Swish(xW) ⊙ (xV) → multiplied through W₂. What did
retracing the chain add? — + a shrunken hidden dimension, and no derivation.
The 2/3 shrink is what makes the comparison fair: without it you’d just be
observing that a bigger FFN does better, which nobody needed a paper for. And
the missing derivation isn’t a gap in this post — it’s the actual state of the
field. The gate won a bake-off, the result replicated in the model families
that adopted it, and that turned out to be enough.
Check yourself
Before you go — a colleague benchmarks SwiGLU against GeLU, keeps the hidden
dimension at the usual 4 · d_model for both, and reports that SwiGLU wins.
Why is that result close to meaningless?
Answer
Because at equal hidden dimension the SwiGLU block has three d_model × d_hidden-ish
matrices where the GeLU block has two — roughly 50% more parameters and more
FLOPs in that sublayer. Of course it wins; you gave it a bigger budget. The
result that made the field switch is the parameter-matched one, which is
exactly why implementations shrink the hidden dimension to about 8/3 · d_model
before comparing. The general move worth keeping: whenever an architecture
change also changes the parameter count, the interesting question is always
“and what did you hold fixed?”
And one more — GeGLU scores about the same as SwiGLU in Shazeer’s own table. So why does essentially every open-weights model ship SwiGLU rather than GeGLU?
Answer
Not because of a decisive result — there isn’t one, and nobody has documented the adoption path directly. The plausible account: when two variants sit within noise of each other, the tiebreaker becomes what the next influential model happened to pick, because that’s what gets copied, what gets a tuned kernel written for it, and what the next paper compares against. PaLM and LLaMA used SwiGLU; the serving stacks fused a SwiGLU kernel; the models after them used SwiGLU. Note that Gemma went the other way and gates with GeLU, which is decent evidence the choice really is close to a coin flip. Treat this as a useful prior for reading architecture choices in general: some are load-bearing, and some are just what the last big model did.
Famous related terms
- Swish —
Swish(x) = x · sigmoid(βx)(oftenβ=1) — a smooth, self-gating activation found via neural architecture search at Google in 2017. Looks like ReLU at the extremes but dips slightly negative around zero, and outperformed ReLU on the architectures the original paper tested. - GeLU —
GeLU(x) = x · Φ(x)— the smooth ReLU that BERT and GPT-2 used. Still the default in many vision transformers. - GLU —
GLU(x) = (xW) ⊙ sigmoid(xV)— the original gated linear unit from Dauphin et al. (2017), built for convolutional language models. SwiGLU is the same shape with Swish swapped in for sigmoid. - GeGLU —
GeGLU(x) = GeLU(xW) ⊙ (xV)— Shazeer’s other top GLU variant. Roughly tied with SwiGLU in Shazeer’s own comparison; SwiGLU is the one that spread, and PaLM and LLaMA picking it is the usual explanation. - Feed-forward layer —
FFN ≈ matmul → activation → matmul— where most of a transformer’s parameters live. The activation function here is what this whole post is about.
Going deeper
- GLU Variants Improve Transformer (Shazeer, 2020) — the primary source: three pages that answer “which gated variants were actually compared, and by how much did they win?”, and end by declining to explain the result.
- Gaussian Error Linear Units (Hendrycks & Gimpel, 2016) — read this for the intermediate step: why a smooth activation with nonzero gradient below zero was worth switching to before gating entered the picture.
- LLaMA paper (Touvron et al., 2023) — the best short answer to “what does this look like in a real model?”, including the 2/3 hidden-dimension adjustment that keeps the parameter count honest.