Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why does softmax look like that?

Softmax is the function that turns a vector of arbitrary numbers into probabilities. The exponential in the middle isn't decorative — it's what makes the whole machine differentiable, well-behaved, and historically inevitable.

Math intermediate Apr 29, 2026 · updated Aug 25, 2026 · 13 min read

On this page

The picture version

Six pictures for a reader who has never seen a model pick a word. The prose below fills in the seams the pictures skip.

1 · The problem

The model produces raw scores. It needs percentages.

“I’m hungry, let’s order ___” what comes out of the model pizza    8.2 sushi    6.0 gravel  −3.1 unbounded, and adding up to nothing in particular one of these per word in the vocabulary pizza    90% sushi    10% gravel 0.001% non-negative, and adding to exactly 100% only now can the model roll dice and choose
The scores are the model’s raw opinion, and they can be any size, positive or negative. Picking a word requires proper percentages, and every token a language model emits passes through whatever does that conversion.

2 · The obvious conversion, and why it breaks

Shift everything positive, then divide by the total. Watch what happens.

shift by just enough 11.3   9.1   0.0 55% / 45% / 0% shift by rather more 108.2   106   96.9 35% / 34% / 31% Same three scores. The answer depends on a number you made up. and shift far enough and gravel becomes almost as likely as pizza
Shifting to remove the negatives and dividing by the sum gives a different answer for every choice of shift — and the choice is arbitrary. Worse, a large shift squeezes all the answers together, so a word the model rated hopeless ends up nearly as likely as its favourite.

3 · The fix

Exponentiate first. Negatives disappear on their own, and gaps become ratios.

any score at all 8.2   6.0   −3.1 exponentiate always positive 3641   403   0.045 divide by the total percentages Nothing was made up. The negative handled itself. And a fixed gap between two scores now always means the same ratio.
Raising a constant to the power of each score makes every value positive without inventing a shift, and turns differences between scores into fixed ratios. Dividing by the total then costs nothing — the hard part was getting positives honestly.

4 · What that buys

Only the gaps matter. Turn every dial up by the same amount and nothing moves.

8.2   6.0   −3.1 90% / 10% / ~0% 1008.2   1006   996.9 90% / 10% / ~0% Identical. Which is precisely what the naive version could not manage.
Adding a thousand to every score changes nothing at all, because the operation only ever sees the differences. The failure of the shift-and-divide approach was that the shift changed the answer; here it cannot — and that property is also what keeps the arithmetic from overflowing in practice.

5 · The property that made it stick

One dial goes up only by taking light from the others.

pizza everyone else raise pizza’s score… …and the rest must shrink The total brightness is fixed at 1, so the options genuinely compete. which also gives training a clean signal: push the right answer up, and the wrong ones fall automatically
Because the shares must total one, raising any option necessarily lowers the others — the candidates are in direct competition rather than scored in isolation. That is what makes the learning signal so well behaved: rewarding the correct word automatically penalises its rivals.

6 · Keep this card

The whole thing on one index card.

softmax = exponentiate every score + divide by the total so only the gaps between scores decide the shares Both halves are forced: the first gets positives honestly, the second makes them add to one.
Picture to keep: a row of adjustable spotlights aimed at a wall, and a rule that the total brightness is always exactly 1. Raising one score doesn’t add light to the room — it steals it from the others. Only the gaps between the scores decide who gets how much, so turning every dial up by the same amount changes nothing.

Why it exists

Picture ChatGPT guessing the next word in “I’m hungry, let’s order ___”. Internally it produces a row of raw scores — say pizza: 8.2, sushi: 6.0, gravel: -3.1, and so on for every word in its vocabulary. Those numbers are unbounded and don’t add up to anything meaningful. Before the model can roll dice and pick one, something has to convert them into actual percentages — for those three scores, pizza 90%, sushi 10%, gravel 0.001%. Softmax is that conversion. Every token a language model emits passes through it.

Anyone who has stared at the softmax formula has had the same small moment of suspicion:

softmax(z)_i = exp(z_i) / Σ_j exp(z_j)

Why exp? Why not just take absolute values, or square, or shift everything to be positive and divide by the sum? Any of those would also turn a vector of arbitrary real numbers into something that sums to 1. The exponential looks like a flourish — a particular choice where many would do.

It isn’t a flourish. The exp is doing several jobs at once that no other elementary function does together, and once you see them, the formula stops looking arbitrary and starts looking like the unique answer to a fairly specific question: what’s the smoothest possible way to turn unbounded “scores” into a probability distribution, in a way a gradient-descent optimizer can actually train?

The neural-network framing is recent. The function itself is older — it’s the same shape as the Boltzmann distribution from 19th-century statistical mechanics, and the same shape as the multinomial logistic regression used in statistics for decades before deep learning showed up. The standard account is that softmax was inherited from those older traditions rather than invented for neural nets — though no clean citation pins down the first paper to put it on top of a neural network classifier, so treat that lineage as “common knowledge in the field” rather than a sourced claim.

Why it matters now

Softmax is how nearly every multi-class classifier and every modern LLM turns its output into probabilities. (Strictly, most frameworks leave it out of the network and fold it into the loss or the sampling step — for numerical reasons covered below — but conceptually it’s the last thing that happens.) When a language model picks the next token, it computes a vector of logits — one real number per vocabulary entry — and softmax converts those into the probability distribution you sample from. The temperature knob you see in chat APIs is literally a divisor inside the softmax. And cross-entropy loss, with a single correct label, reduces to −log of the softmax output for that class.

If softmax is wrong, training is wrong. If softmax is numerically unstable, your loss explodes. If you swap it for something “simpler” without understanding what exp was doing, you usually discover the model stops learning. It’s load-bearing in a quiet way.

The short answer

softmax = exponentiate every score + divide by the total, i.e. softmax(z)_i = exp(z_i) / Σ_j exp(z_j)

Picture to keep: a row of adjustable spotlights aimed at a wall, and a rule that the total brightness is always exactly 1. Raising one score doesn’t add light to the room — it steals it from the others. Only the gaps between the scores decide who gets how much; turning every dial up by the same amount changes nothing. (Where the picture breaks: the brightness isn’t conserved by any physical law, and the trade isn’t linear — one dial pulled far enough ahead takes essentially all the light, because the exponential is doing the sharing.)

It’s the function that takes a vector of real numbers (positive, negative, unbounded) and returns a probability distribution where each entry is proportional to exp of its score. The exp makes everything positive, makes the ratios depend only on differences between scores, and makes the gradients clean. There isn’t a simpler function that does all three.

How it works

The naive attempt, and why it breaks

Take the three scores from the hook — pizza: 8.2, sushi: 6.0, gravel: −3.1 — and do the obvious thing. Shift everything up until nothing is negative, then divide by the sum. Add 3.1 to each:

11.3, 9.1, 0.0   →   sum 20.4   →   pizza 55%, sushi 45%, gravel 0%

Two things are wrong, and both are fatal.

Gravel gets exactly zero. Not “vanishingly small” — zero. A probability of exactly 0 has zero gradient, so if the training data ever does say gravel, the model has no slope to climb back up. And −log 0 is infinite, which is what your loss will print.

The answer depends on a number you made up. Shift by 100 instead of 3.1 — same scores, same gaps, an equally defensible choice:

108.2, 106.0, 96.9   →   sum 311.1   →   pizza 35%, sushi 34%, gravel 31%

Pizza went from a 55% favorite to a coin flip against gravel, and nothing about the model’s opinion changed. The shift-and-divide scheme lets an arbitrary constant rewrite the model’s beliefs. That’s not a rough edge; it means the function is measuring the wrong thing. What we want it to measure is the gaps between scores, and shift-and-divide measures the ratios of the shifted values instead.

The fix: four constraints, and only one function survives

So write down what we actually need, and watch how little freedom is left.

  1. Output is a probability distribution. All entries non-negative, sum to 1.
  2. Strictly increasing in each score. Raising z_i should raise p_i, never lower it.
  3. Translation invariant. Adding the same constant to every score shouldn’t change the output. (If all scores go up by 5, no class became more likely relative to the others.)
  4. Smooth and differentiable everywhere. This is the gradient-descent constraint. We want to backpropagate through it.

The first constraint suggests dividing by the sum: p_i = f(z_i) / Σ_j f(z_j) for some non-negative f. The second says f must be increasing. The third — translation invariance — is the surprisingly strong one. Translation invariance says f(z_i + c) / f(z_j + c) must equal f(z_i) / f(z_j) for all c. The only continuous function satisfying that multiplicative property under addition is the exponential: f(z) = exp(α·z) for some α > 0. (This is essentially the “Cauchy functional equation” in disguise.)

So the exp isn’t a stylistic pick — it’s forced by wanting the function to depend only on score differences, which is what we mean by “the absolute scale of logits is meaningless, only their gaps matter.” The free parameter α is exactly the inverse temperature: softmax_T(z)_i = exp(z_i / T) / Σ exp(z_j / T).

Run it on the hook’s scores and you get pizza 90%, sushi 10%, gravel 0.001% — and, crucially, you get the same answer whether you feed it 8.2, 6.0, −3.1 or 108.2, 106.0, 96.9. Gravel is tiny but never exactly zero, so the gradient always has somewhere to go.

Why translation invariance matters in practice

This is the part where the seam shows. Logits coming out of a neural network can be enormous — exp(1000) overflows a 32-bit float instantly. But because softmax is translation invariant, you can subtract max(z) from every entry before exponentiating without changing the answer. The largest input becomes 0, exp(0) = 1, everything else is between 0 and 1, no overflow. Every production softmax implementation does this. It’s a free numerical-stability trick that falls out of the same property that justified the exp in the first place.

The gradient that made it stick

Even if you accepted some other positive-and-increasing function, softmax has one more property that explains why the field standardized on it: when you compose softmax with cross-entropy loss, the gradient simplifies to something almost embarrassingly clean.

If p = softmax(z) and the loss is L = −log p_y (cross-entropy with the correct class y), then:

∂L/∂z_i = p_i − 1[i = y]

That’s it. The gradient with respect to the logits is just “predicted probability minus 1 for the right class.” No exponentials in the gradient, no chain-rule explosion, no special cases. Every modern deep-learning framework exploits this by fusing softmax and cross-entropy into a single op (log_softmax + nll_loss, or softmax_cross_entropy_with_logits) because computing them separately is both slower and less numerically stable.

That clean gradient is, I’d argue, the single biggest reason softmax survived the transition from statistics to deep learning. The four constraints already pin the function down to exp(α·z) for some α > 0 — that is, to softmax at some temperature. The gradient is what made the field stop fiddling with α and treat the plain version as the default.

The Boltzmann connection

The same formula p_i ∝ exp(−E_i / kT) describes the probability of finding a thermodynamic system in state i with energy E_i at temperature T. The formulas line up term for term, up to scaling: a logit plays the role of a negative energy, and the softmax temperature plays the role of the physical temperature. High temperature → the distribution flattens (all states roughly equally likely). Low temperature → it concentrates on the lowest-energy (highest-logit) state. At T → 0, softmax becomes argmax — or, if two logits are exactly tied, a uniform split over the tied winners.

The standard account in physics derives the Boltzmann distribution by maximizing entropy subject to a constraint on expected energy. Run the same derivation in ML — what’s the maximum-entropy distribution consistent with these expected sufficient statistics? — and you get softmax. Same math, different vocabulary.

The physics analogy is exact on the formula and misleading everywhere else: there’s no energy, no heat, and no equilibrium in a language model. A logit is not a physical energy; it’s a learned number that happens to sit in the same slot in the same equation.

You started with softmax(z)_i = exp(z_i) / Σ_j exp(z_j). What did this post add? — + exp is not a choice. Demand that only the gaps between scores matter and the exponential is the only continuous function left standing. Everything else people like about softmax — the max-subtraction stability trick, the one-line gradient, the temperature knob — is a consequence of that one demand, not a separate feature.

Check yourself

Before you go — a colleague “simplifies” softmax to p_i = z_i² / Σ_j z_j², noting that squaring also makes everything positive and the outputs still sum to 1. Which of the four constraints does it break, and what does that do to training?

Answer

Squaring isn’t monotonic in z. A logit of −5 and a logit of +5 produce the same probability, so the model can lower a score and raise its probability — constraint 2 is violated, and gradient descent gets a signal pointing the wrong way on half the number line. A logit at exactly 0 also gets probability 0 with zero gradient, the same trap as shift-and-divide. And because the scheme isn’t translation invariant either, adding a constant to every score changes the answer. Squaring fixes the sign and nothing else.

And one more — you set temperature to 0.01 on a model whose top two logits differ by 3.0. Roughly what does the output distribution look like, and what does that do to sampling?

Answer

Dividing by T = 0.01 multiplies every gap by 100, so a gap of 3.0 becomes 300. exp(300) against exp(0) means the top token takes essentially all the probability mass — the distribution collapses to a spike. Sampling from it is indistinguishable from taking the argmax, which is exactly the T → 0 limit named in the post. Practically, output becomes deterministic and repetitive, and the implementation is only saved from overflow by the max-subtraction trick, since exp(300) would blow up a float32 on its own.

Going deeper

A note on what I’m sure of. The translation-invariance argument forcing exp is a standard piece of the exponential family story and the clean softmax–cross-entropy gradient is straightforward calculus. The historical claim about who first used softmax on a neural network is the part I’d most want a real source for and don’t have one on hand; the broader “this came from statistical mechanics and multinomial logistic regression” framing is the standard account but I haven’t traced individual citations.