Why does softmax look like that?
Softmax is the function that turns a vector of arbitrary numbers into probabilities. The exponential in the middle isn't decorative — it's what makes the whole machine differentiable, well-behaved, and historically inevitable.
On this page
- The picture version
- Why it exists
- Why it matters now
- The short answer
- How it works
- The naive attempt, and why it breaks
- The fix: four constraints, and only one function survives
- Why translation invariance matters in practice
- The gradient that made it stick
- The Boltzmann connection
- Check yourself
- Famous related terms
- Going deeper
The picture version
Six pictures for a reader who has never seen a model pick a word. The prose below fills in the seams the pictures skip.
1 · The problem
The model produces raw scores. It needs percentages.
2 · The obvious conversion, and why it breaks
Shift everything positive, then divide by the total. Watch what happens.
3 · The fix
Exponentiate first. Negatives disappear on their own, and gaps become ratios.
4 · What that buys
Only the gaps matter. Turn every dial up by the same amount and nothing moves.
5 · The property that made it stick
One dial goes up only by taking light from the others.
6 · Keep this card
The whole thing on one index card.
Why it exists
Picture ChatGPT guessing the next word in “I’m hungry, let’s order ___”.
Internally it produces a row of raw scores — say pizza: 8.2, sushi: 6.0,
gravel: -3.1, and so on for every word in its vocabulary. Those numbers
are unbounded and don’t add up to anything meaningful. Before the model can
roll dice and pick one, something has to convert them into actual
percentages — for those three scores, pizza 90%, sushi 10%, gravel 0.001%.
Softmax is that conversion. Every token a language model emits passes
through it.
Anyone who has stared at the softmax formula has had the same small moment of suspicion:
softmax(z)_i = exp(z_i) / Σ_j exp(z_j)
Why exp? Why not just take absolute values, or square, or shift everything
to be positive and divide by the sum? Any of those would also turn a vector
of arbitrary real numbers into something that sums to 1. The exponential
looks like a flourish — a particular choice where many would do.
It isn’t a flourish. The exp is doing several jobs at once that no other
elementary function does together, and once you see them, the formula stops
looking arbitrary and starts looking like the unique answer to a fairly
specific question: what’s the smoothest possible way to turn unbounded
“scores” into a probability distribution, in a way a gradient-descent
optimizer can actually train?
The neural-network framing is recent. The function itself is older — it’s the same shape as the Boltzmann distribution from 19th-century statistical mechanics, and the same shape as the multinomial logistic regression used in statistics for decades before deep learning showed up. The standard account is that softmax was inherited from those older traditions rather than invented for neural nets — though no clean citation pins down the first paper to put it on top of a neural network classifier, so treat that lineage as “common knowledge in the field” rather than a sourced claim.
Why it matters now
Softmax is how nearly every multi-class classifier and every modern
LLM
turns its output into probabilities. (Strictly, most frameworks leave it out
of the network and fold it into the loss or the sampling step — for numerical
reasons covered below — but conceptually it’s the last thing that happens.)
When a language model picks the next token, it computes a vector of
logits
— one real number per vocabulary entry — and softmax converts those into
the probability distribution you sample from. The temperature knob you
see in chat APIs is literally a divisor inside the softmax. And cross-entropy
loss, with a single correct label, reduces to −log of the softmax output for
that class.
If softmax is wrong, training is wrong. If softmax is numerically unstable,
your loss explodes. If you swap it for something “simpler” without
understanding what exp was doing, you usually discover the model stops
learning. It’s load-bearing in a quiet way.
The short answer
softmax = exponentiate every score + divide by the total, i.e. softmax(z)_i = exp(z_i) / Σ_j exp(z_j)
Picture to keep: a row of adjustable spotlights aimed at a wall, and a rule that the total brightness is always exactly 1. Raising one score doesn’t add light to the room — it steals it from the others. Only the gaps between the scores decide who gets how much; turning every dial up by the same amount changes nothing. (Where the picture breaks: the brightness isn’t conserved by any physical law, and the trade isn’t linear — one dial pulled far enough ahead takes essentially all the light, because the exponential is doing the sharing.)
It’s the function that takes a vector of real numbers (positive, negative,
unbounded) and returns a probability distribution where each entry is
proportional to exp of its score. The exp makes everything positive,
makes the ratios depend only on differences between scores, and makes the
gradients clean. There isn’t a simpler function that does all three.
How it works
The naive attempt, and why it breaks
Take the three scores from the hook — pizza: 8.2, sushi: 6.0,
gravel: −3.1 — and do the obvious thing. Shift everything up until nothing
is negative, then divide by the sum. Add 3.1 to each:
11.3, 9.1, 0.0 → sum 20.4 → pizza 55%, sushi 45%, gravel 0%
Two things are wrong, and both are fatal.
Gravel gets exactly zero. Not “vanishingly small” — zero. A probability of
exactly 0 has zero gradient, so if the training data ever does say gravel,
the model has no slope to climb back up. And −log 0 is infinite, which is
what your loss will print.
The answer depends on a number you made up. Shift by 100 instead of 3.1 — same scores, same gaps, an equally defensible choice:
108.2, 106.0, 96.9 → sum 311.1 → pizza 35%, sushi 34%, gravel 31%
Pizza went from a 55% favorite to a coin flip against gravel, and nothing about the model’s opinion changed. The shift-and-divide scheme lets an arbitrary constant rewrite the model’s beliefs. That’s not a rough edge; it means the function is measuring the wrong thing. What we want it to measure is the gaps between scores, and shift-and-divide measures the ratios of the shifted values instead.
The fix: four constraints, and only one function survives
So write down what we actually need, and watch how little freedom is left.
- Output is a probability distribution. All entries non-negative, sum to 1.
- Strictly increasing in each score. Raising
z_ishould raisep_i, never lower it. - Translation invariant. Adding the same constant to every score shouldn’t change the output. (If all scores go up by 5, no class became more likely relative to the others.)
- Smooth and differentiable everywhere. This is the gradient-descent constraint. We want to backpropagate through it.
The first constraint suggests dividing by the sum: p_i = f(z_i) / Σ_j f(z_j) for some non-negative f. The second says f must be increasing.
The third — translation invariance — is the surprisingly strong one.
Translation invariance says f(z_i + c) / f(z_j + c) must equal f(z_i) / f(z_j) for all c. The only continuous function satisfying that
multiplicative property under addition is the exponential: f(z) = exp(α·z) for some α > 0. (This is essentially the “Cauchy functional
equation” in disguise.)
So the exp isn’t a stylistic pick — it’s forced by wanting the function
to depend only on score differences, which is what we mean by “the
absolute scale of logits is meaningless, only their gaps matter.” The free
parameter α is exactly the inverse temperature: softmax_T(z)_i = exp(z_i / T) / Σ exp(z_j / T).
Run it on the hook’s scores and you get pizza 90%, sushi 10%, gravel 0.001%
— and, crucially, you get the same answer whether you feed it 8.2, 6.0, −3.1 or 108.2, 106.0, 96.9. Gravel is tiny but never exactly zero, so the
gradient always has somewhere to go.
Why translation invariance matters in practice
This is the part where the seam shows. Logits coming out of a neural
network can be enormous — exp(1000) overflows a 32-bit float
instantly. But because softmax is translation invariant, you can subtract
max(z) from every entry before exponentiating without changing the
answer. The largest input becomes 0, exp(0) = 1, everything else is
between 0 and 1, no overflow. Every production softmax implementation does
this. It’s a free numerical-stability trick that falls out of the same
property that justified the exp in the first place.
The gradient that made it stick
Even if you accepted some other positive-and-increasing function, softmax has one more property that explains why the field standardized on it: when you compose softmax with cross-entropy loss, the gradient simplifies to something almost embarrassingly clean.
If p = softmax(z) and the loss is L = −log p_y (cross-entropy with the
correct class y), then:
∂L/∂z_i = p_i − 1[i = y]
That’s it. The gradient with respect to the logits is just “predicted
probability minus 1 for the right class.” No exponentials in the gradient,
no chain-rule explosion, no special cases. Every modern deep-learning
framework exploits this by fusing softmax and cross-entropy into a single
op (log_softmax + nll_loss, or softmax_cross_entropy_with_logits)
because computing them separately is both slower and less numerically
stable.
That clean gradient is, I’d argue, the single biggest reason softmax
survived the transition from statistics to deep learning. The four constraints
already pin the function down to exp(α·z) for some α > 0 — that is, to
softmax at some temperature. The gradient is what made the field stop
fiddling with α and treat the plain version as the default.
The Boltzmann connection
The same formula p_i ∝ exp(−E_i / kT) describes the probability of
finding a thermodynamic system in state i with energy E_i at
temperature T. The formulas line up term for term, up to scaling: a logit
plays the role of a negative energy, and the softmax temperature plays the
role of the physical temperature. High temperature →
the distribution flattens (all states roughly equally likely). Low
temperature → it concentrates on the lowest-energy (highest-logit) state.
At T → 0, softmax becomes argmax — or, if two logits are exactly tied,
a uniform split over the tied winners.
The standard account in physics derives the Boltzmann distribution by maximizing entropy subject to a constraint on expected energy. Run the same derivation in ML — what’s the maximum-entropy distribution consistent with these expected sufficient statistics? — and you get softmax. Same math, different vocabulary.
The physics analogy is exact on the formula and misleading everywhere else: there’s no energy, no heat, and no equilibrium in a language model. A logit is not a physical energy; it’s a learned number that happens to sit in the same slot in the same equation.
You started with softmax(z)_i = exp(z_i) / Σ_j exp(z_j). What did this post
add? — + exp is not a choice. Demand that only the gaps between scores
matter and the exponential is the only continuous function left standing.
Everything else people like about softmax — the max-subtraction stability
trick, the one-line gradient, the temperature knob — is a consequence of that
one demand, not a separate feature.
Check yourself
Before you go — a colleague “simplifies” softmax to p_i = z_i² / Σ_j z_j²,
noting that squaring also makes everything positive and the outputs still sum
to 1. Which of the four constraints does it break, and what does that do to
training?
Answer
Squaring isn’t monotonic in z. A logit of −5 and a logit of +5 produce
the same probability, so the model can lower a score and raise its
probability — constraint 2 is violated, and gradient descent gets a signal
pointing the wrong way on half the number line. A logit at exactly 0 also gets
probability 0 with zero gradient, the same trap as shift-and-divide. And
because the scheme isn’t translation invariant either, adding a constant to
every score changes the answer. Squaring fixes the sign and nothing else.
And one more — you set temperature to 0.01 on a model whose top two logits differ by 3.0. Roughly what does the output distribution look like, and what does that do to sampling?
Answer
Dividing by T = 0.01 multiplies every gap by 100, so a gap of 3.0 becomes
300. exp(300) against exp(0) means the top token takes essentially all the
probability mass — the distribution collapses to a spike. Sampling from it is
indistinguishable from taking the argmax, which is exactly the T → 0 limit
named in the post. Practically, output becomes deterministic and repetitive,
and the implementation is only saved from overflow by the max-subtraction
trick, since exp(300) would blow up a float32 on its own.
Famous related terms
- Sigmoid —
sigmoid(x) = 1 / (1 + exp(−x))— softmax for two classes. Falls out of softmax over{x, 0}after cancellation. Same exponential, same translation-invariance argument. - Logit —
logit(p) = log(p / (1 − p))— the inverse of sigmoid. The “raw scores” that go into softmax are called logits because in the binary case they literally are logits. - Cross-entropy loss —
H(p, q) = − Σ p(x) log q(x)— the loss function softmax was made for; see entropy. - Temperature —
softmax_T(z) = softmax(z / T)— the knob that controls how peaked or flat the output distribution is. See temperature. - Boltzmann distribution —
p_i ∝ exp(−E_i / kT)— the physics ancestor. Maximum-entropy distribution given an energy constraint. - Argmax —
argmax = softmax at T → 0— the non-differentiable function softmax exists to smooth out.
Going deeper
- Goodfellow, Bengio, Courville, Deep Learning
— for “where does the clean
p − ygradient come from, and why is softmax fused into the loss?”, derived in the chapter on output units. I’d treat this as the closest thing to a canonical reference here: there’s no single origin paper I can point at with confidence. - Bridle, J. S. (1990), Probabilistic interpretation of feedforward classification network outputs — for “who put softmax on a neural net first?”, one of the early treatments and the usual citation, though the attribution is conventional rather than settled — treat it as a search starting point.
- Rabbit hole: Jaynes, Probability Theory: The Logic of Science — for “why does the same formula show up in thermodynamics?”, answered as constrained entropy maximization rather than as physics.
A note on what I’m sure of. The translation-invariance argument forcing
expis a standard piece of the exponential family story and the clean softmax–cross-entropy gradient is straightforward calculus. The historical claim about who first used softmax on a neural network is the part I’d most want a real source for and don’t have one on hand; the broader “this came from statistical mechanics and multinomial logistic regression” framing is the standard account but I haven’t traced individual citations.