Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why SwiGLU replaced ReLU in transformers

Modern LLMs ditched the simplest activation function in deep learning for a multiplicative gate nobody can fully explain. Here's why.

AI & ML intermediate Apr 29, 2026 · updated Aug 25, 2026 · 10 min read

On this page

The picture version

Five pictures for a reader who has never opened a transformer’s source. The prose below fills in the seams the pictures skip.

1 · The problem

The textbook promised two matrices. The code has three.

what every diagram shows project up a nonlinearity project down two matrices, one nonlinearity what the source actually has gate_proj up_proj × elementwise down_proj One branch multiplies the other. No textbook mentioned that.
That block — one feed-forward layer of one modern model — is the running example for this post. You’d assume a change like this came from theory. It’s closer to the opposite: the three-matrix block won an empirical bake-off, and the paper that established the result declined to explain why.

2 · Complaint one

Below zero, ReLU stops teaching.

exactly flat gradient = 0 ReLU dips below, then returns — still a signal just under zero GeLU A neuron stuck in the flat region never learns its way out. the fix is to the flat region near zero, not to the far tail — BERT and GPT-2 picked GeLU up and it became the new default
The curves are the shape of each function, not plotted values. The empirical gains were small but consistent, which turns out to be the pattern for this whole lineage — and sets up why the next step is so hard to justify from theory.

3 · The move

Let one branch decide how much of the other gets through.

x a projection, then Swish this branch is the gate a plain projection this branch is the content × project back down A learned, per-element volume knob on the block’s own output. the hidden width is cut to two-thirds so the three matrices cost about what two used to — the comparison is held at equal parameters
Swap Swish for GeLU on the gate branch and you get GeGLU instead — the same shape, a different activation. Google’s Gemma line does exactly that, which is a useful reminder that the shape is the durable part and the activation on it is not.

4 · The seam

The paper that settled it declined to say why it works.

the variants won the bake-off and the paper offers no account of why So what settled it wasn’t theory. It was what the next big model shipped. when two variants sit within noise of each other, the tiebreaker becomes what gets copied, what gets a tuned kernel written for it, and what the next paper compares against that account is plausible rather than documented — nobody has traced the adoption path directly and Gemma going the other way is decent evidence the choice really is close to a coin flip
What is on the record is narrower than “nobody knows”: the author explicitly declined to explain the result, and the field has no settled mechanistic account. Useful prior for reading architecture choices in general — some are load-bearing, and some are just what the last big model did.

5 · Keep this card

The whole thing on one index card.

SwiGLU = a feed-forward block + a second projection used as a learned gate ∴ adopted because it measured better, not because anyone derived it
Picture to keep: two wires leaving the same input — one carrying the signal, one carrying a per-element volume knob that was itself learned. Multiply them together and you get an FFN that can suppress its own output channel by channel. The third matrix in your Llama layer is that second wire.

Why it exists

Open the source of any recent open-weights model — Llama, Qwen, Mistral — and scroll to the FFN block inside one transformer layer. That’s our running example for this post: one FFN block, in one layer, of one modern model. You expect two weight matrices, because that’s what every textbook and every diagram of a transformer shows: project up, apply a nonlinearity, project back down. Instead you find three — gate_proj, up_proj, down_proj — and an elementwise multiplication sitting between them that no textbook mentioned.

You’d assume a change like that came from theory: someone proved the old activation was leaving something on the table. It’s closer to the opposite. The three-matrix block won an empirical bake-off, and the author of the paper that established the result explicitly declined to explain why it works. The interesting question isn’t “what is SwiGLU” — it’s how a field ends up standardizing on something nobody can derive.

If you opened a deep-learning textbook in 2016, the activation function on every page was ReLU. It is the simplest nonlinearity that works: a kink at zero, a straight line above it, no exponentials. ResNets used it. The original Transformer paper used it. For a long time it was the default for the same reason printf is the default debug tool — it’s cheap, it’s understood, it gets out of the way.

Then something quietly happened between 2017 and 2022. BERT and GPT-2 swapped ReLU for GeLU. Then PaLM and LLaMA swapped GeLU for SwiGLU. By 2024, opening most decoder-only open-weights models, you’d find a feed-forward layer with three matrices instead of two and a multiplication you’d never seen in a textbook. Something pushed a large part of the field to abandon a thing that was famously fine.

Why it matters now

Most modern LLM serving stacks — vLLM and TensorRT-LLM among them — ship a fused kernel for the gated feed-forward block, because the three-matrix shape is what the Llama, Qwen, and Mistral families all use. It isn’t universal: Google’s Gemma line, for instance, gates with GeLU rather than Swish — same three-matrix shape, different activation on the gate branch. Quantization schemes and tensor-parallel splits have to handle that shape specifically. If you’re reading LLaMA or Qwen or Mistral source code and the mlp block has gate_proj, up_proj, and down_proj, that’s a gated FFN — in those models specifically, SwiGLU. Knowing why that shape won is the difference between “the FFN is a black box” and “I can predict how this kernel allocates memory.”

The short answer

SwiGLU = Swish(xW) ⊙ (xV) → multiplied through W₂

Picture to keep: a mixing desk. One branch carries the signal; the other branch is a row of faders, one per channel, that the network learns to set — and the two get multiplied together before anything leaves the block. ReLU only had a single on/off switch per channel; SwiGLU has a continuous fader whose position depends on the input. (Where the analogy breaks: on a real mixing desk a human sets the faders once; here both branches are computed fresh from the same input on every token, so “signal” and “faders” aren’t fixed roles — the split is a convention of the formula, not a fact about which branch means what.)

In words: instead of the classic feed-forward layer f(xW₁) · W₂, you compute two projections of the input, pass one through a smooth activation (Swish), multiply them elementwise (the “gate”), and then project back down with a third matrix W₂. The gate lets the network decide, per coordinate, how much of the other branch’s signal to let through.

How it works

The three matrices in your Llama layer are the residue of a chain of complaints. Each step below is somebody’s fix for the previous step’s annoyance — which is why you can’t derive the shape from first principles, only retrace it.

The classic FFN block (Vaswani et al., 2017). Given input x, compute h = ReLU(x W₁ + b₁), then y = h W₂ + b₂. Two matmuls, one nonlinearity. The hidden dimension is typically 4× the model dimension — that’s where most of the parameter count of a transformer actually lives, more than attention.

What breaks in the classic block: ReLU has two annoyances.

Step 1: ReLU → GeLU (around 2018). Below zero it is exactly flat — the gradient is zero, so a neuron stuck in the negative region gets no learning signal (“dying ReLU”). And the kink at zero means the function isn’t differentiable there, which is fine in practice but ugly in theory. Hendrycks & Gimpel’s GeLU paper (arXiv:1606.08415) proposed x · Φ(x) where Φ is the Gaussian CDF. Same shape as ReLU at the extremes, but smooth, and crucially its gradient isn’t identically zero for negative inputs, so a neuron sitting just below zero still gets a learning signal. (It does still decay toward zero far out on the negative side — the fix is to the flat region near zero, not to the tail.) BERT and GPT-2 picked it up and it became the new default. The empirical gains were small but consistent.

Step 2: GeLU → SwiGLU (around 2020-2022). This is the weirder jump. Noam Shazeer’s “GLU Variants Improve Transformer” (arXiv:2002.05202, 2020) tested a family of GLU variants in the FFN sublayer. The pattern is:

SwiGLU(x) = (Swish(x W) ⊙ (x V)) W₂

Two input projections (W and V) instead of one. One goes through Swish, the other stays linear. They get multiplied elementwise (⊙). Then W₂ projects back. The gate is the new ingredient — the linear branch can amplify or suppress each coordinate of the activated branch.

Why “free performance”? To keep the parameter count comparable to the classic two-matrix FFN, SwiGLU implementations shrink the hidden dimension by 2/3 (so roughly 8/3 · d_model instead of 4 · d_model). With matched parameters, SwiGLU still outperforms GeLU and ReLU on perplexity and downstream benchmarks. PaLM used it, LLaMA 1/2/3 use it, and most open-weights models published after late 2022 use it.

Why does the gate help? Honestly — and this is the seam — nobody really knows. Shazeer’s paper closes with a now-famous line attributing the improvement to “divine benevolence.” The hand-wavy story is that multiplicative interactions let the network represent things linear-plus-ReLU can’t easily represent (e.g. quadratic functions of the input), and the gate gives it a learnable per-coordinate dial. But the honest answer is that we have an empirical result that replicates across scales, and a mechanism story that’s plausible but not proven. The field adopted it because the loss curves were better, not because someone derived it from first principles.

One more thing — the cost. SwiGLU has three big matmuls in the FFN (gate_proj, up_proj, down_proj) instead of two. That changes how kernels fuse, how TP splits the weights, and how quantization schemes group the channels. If you’ve ever wondered why the major inference engines ship a dedicated fused kernel for this block, that’s why — and it’s why the block in your Llama layer has three names in it instead of two.

You started with SwiGLU = Swish(xW) ⊙ (xV) → multiplied through W₂. What did retracing the chain add? — + a shrunken hidden dimension, and no derivation. The 2/3 shrink is what makes the comparison fair: without it you’d just be observing that a bigger FFN does better, which nobody needed a paper for. And the missing derivation isn’t a gap in this post — it’s the actual state of the field. The gate won a bake-off, the result replicated in the model families that adopted it, and that turned out to be enough.

Check yourself

Before you go — a colleague benchmarks SwiGLU against GeLU, keeps the hidden dimension at the usual 4 · d_model for both, and reports that SwiGLU wins. Why is that result close to meaningless?

Answer

Because at equal hidden dimension the SwiGLU block has three d_model × d_hidden-ish matrices where the GeLU block has two — roughly 50% more parameters and more FLOPs in that sublayer. Of course it wins; you gave it a bigger budget. The result that made the field switch is the parameter-matched one, which is exactly why implementations shrink the hidden dimension to about 8/3 · d_model before comparing. The general move worth keeping: whenever an architecture change also changes the parameter count, the interesting question is always “and what did you hold fixed?”

And one more — GeGLU scores about the same as SwiGLU in Shazeer’s own table. So why does essentially every open-weights model ship SwiGLU rather than GeGLU?

Answer

Not because of a decisive result — there isn’t one, and nobody has documented the adoption path directly. The plausible account: when two variants sit within noise of each other, the tiebreaker becomes what the next influential model happened to pick, because that’s what gets copied, what gets a tuned kernel written for it, and what the next paper compares against. PaLM and LLaMA used SwiGLU; the serving stacks fused a SwiGLU kernel; the models after them used SwiGLU. Note that Gemma went the other way and gates with GeLU, which is decent evidence the choice really is close to a coin flip. Treat this as a useful prior for reading architecture choices in general: some are load-bearing, and some are just what the last big model did.

Going deeper