Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why beam search died for LLMs

Beam search was the default way to decode neural sequence models for years. Then chatbots arrived and quietly stopped using it. The reason is stranger than 'sampling is more creative.'

AI & ML intermediate Apr 29, 2026 · updated Aug 25, 2026 · 11 min read

On this page

The picture version

Six pictures for a reader who has never chosen a decoding strategy. The prose below fills in the seams the pictures skip.

1 · The problem

One box repeats itself on purpose. The other refuses to.

translate, twice “the cat sat on the mat” “the cat sat on the mat” identical regenerate, twice “In a hidden valley, unicorns…” “Nobody expected the herd…” different every time Same kind of machine. Opposite decision about how to walk its probabilities. both are neural nets predicting one token at a time — the difference is entirely in the decoder bolted on afterwards
The left-hand box uses beam search: hunt for the sequence the model scores highest. The right-hand box samples. This post is about why the second one had to abandon the first one’s method.

2 · What beam search is

Keep the best few half-sentences alive instead of committing to one.

step 1 step 2 — expand each, score all, keep the best 2 again The A The unicorns -1.2 ✓ The valley -3.7 A herd -2.9 A single -2.0 ✓ the two best survive to step 3 the rest are dropped, however good they looked one token ago Wider beam = more search = a higher-probability sentence found. Always. with a beam of 1 this is plain greedy decoding; scores are illustrative cumulative log-probabilities
Greedy decoding commits to the best token now and can be trapped by it; beam search keeps k candidates alive so a locally worse token can still win if it leads somewhere better. The search half of beam search is not what fails — it does exactly what it promises.

3 · The surprise

Search harder and the writing gets worse.

maximise — the model’s most likely text “The unicorns were extremely friendly. The unicorns were extremely friendly. The unicorns were extremely friendly…” paraphrased, not a quote from the paper sample — a draw from the same distribution “The herd kept its distance at first, then one of them stepped forward and looked at us for a long moment.” ordinary prose Same model. Same numbers. The loop is genuinely scored higher. so this is not a bug in the search — a better search would find the loop faster
Holtzman et al. (2019) documented the effect and drew the load-bearing conclusion: maximisation is an inappropriate decoding objective for open-ended generation. The fix is not a smarter search; it is a different objective. Why the distribution has this shape is still debated.

4 · Why

The peak of the landscape is not where the good writing lives.

every paragraph the model could possibly write about unicorns probability the argmax one looping sentence what beam search returns the broad plateau of perfectly good paragraphs what sampling draws from “Most probable” and “good” are two different objects.
Shape of the argument, not measured data. When one right answer exists — a translation, a transcription — the model piles probability near it and the peak is the answer. When a million valid paragraphs exist, the mass spreads out and the peak becomes a degenerate outlier nobody wanted.

5 · Even on home turf

Translation has the same disease, in a milder form.

beam width → translation quality the sweet spot a modest beam, not a huge one past the sweet spot, wider beams make it worse what the wider beam finds higher-probability sequences that are systematically too short length normalisation pushes the peak wider, but very large beams still hurt The search is fine. The score is what you should not have trusted.
Known as the beam search curse. Sequence probability is a product of per-token probabilities, so shorter sequences pay fewer multiplicative penalties — a bias toward brevity that a weak search was accidentally hiding. Curve is the shape of the reported effect, not plotted data.

6 · Keep this card

Which half of the definition died.

beam search = search for the most probable sentence + the assumption that most probable means best ∴ the search survived. the assumption is what died. Google Translate keeps it, because there really is roughly one right answer
Picture to keep: a crowd of a million people who each wrote one paragraph about unicorns. Sampling hands you a random paragraph from the pile; beam search hunts for the single paragraph the crowd was most likely to have produced — and that turns out to be one dull sentence copy-pasted six times.

Why it exists

Paste a sentence into Google Translate, hit translate twice, and you get the same output both times. Ask a chatbot to write you a paragraph about unicorns, hit regenerate, and you get a different paragraph. Same underlying kind of model — a neural net predicting one token at a time — and yet one behaves like a lookup and the other behaves like a die roll. That difference isn’t an accident of branding. It’s a decision about how to walk the model’s probabilities, and the two systems make opposite choices. Keep both cases in mind; the whole post is about why the second one had to abandon the first one’s method.

The first one’s method is beam search. Through the mid-2010s it was the workhorse decoder for NMT, summarization, and image captioning. The intuition is clean: greedy decoding picks the locally best token and gets stuck; beam search keeps the top-k partial sequences alive at every step, so it has a shot at finding a globally higher-probability sequence. More search, better answer.

Then LLMs arrived, ChatGPT shipped, and beam search quietly stopped being the default for open-ended generation. The OpenAI chat completions API doesn’t expose it; what it exposes are sampling knobs — temperature and top-p — and no beam width.

That’s the puzzle. You probably assume beam search fell out of favour because sampling is “more creative” — a deliberate trade of quality for variety. It’s very nearly the opposite: on open-ended prompts, the harder you search for the model’s best answer, the worse the text you get.

Why it matters now

Every time you call a chat model, something is choosing one of tens of thousands of tokens per step. That choice is a strategy, not a fact about the model. Picking the wrong strategy makes a state-of-the-art model produce slop — bland paragraphs, repetitive loops, weirdly short answers. Engineers who reach for beam search expecting “higher quality output” get the opposite, and the failure mode looks like the model itself is bad.

The shift also reshaped how people think about prompting. If decoding is sampling, then “the answer” isn’t a thing the model has — it’s a distribution the model has, and you’re drawing from it. That mental model is load-bearing for understanding why the same prompt gives different outputs, why temperature matters, and why two perfectly correct answers can sit side by side.

The short answer

beam search = search for the most probable sentence + the assumption that "most probable" means "best"

Picture to keep: a crowd of a million people who each wrote one paragraph about unicorns. Sampling hands you a random paragraph from the pile. Beam search hunts for the single paragraph the crowd was most likely to have produced — and that turns out to be one dull sentence copy-pasted six times. The analogy breaks in one place worth naming: the model isn’t averaging anyone’s paragraphs. It’s assigning a probability to every possible paragraph, and the peak of that landscape simply isn’t where the good writing lives.

Beam search optimizes for the highest-probability sequence under the model. That assumption holds when there is one right answer and the model concentrates its probability near it — translating your sentence, transcribing speech. It fails when there are many valid continuations, because then the highest-probability sequence is a degenerate one nobody wanted. For open-ended generation, the most probable output is worse than a randomly sampled one. So sampling won.

How it works

The mechanism is counter-intuitive enough that it’s worth walking through.

Beam search, briefly

At each step, keep the top-k partial sequences (the “beam”) ranked by cumulative log-probability. Expand each by one token, score all kN candidates, keep the top k again. At the end, return the highest-scoring full sequence. With k=1 you get greedy decoding; as k grows, you approach exact MAP decoding (true exhaustive argmax over sequences requires more than just a wide beam, but the intuition is right).

For a translation system, this is great. There’s roughly one correct translation; the model concentrates probability on tokens near that translation; searching harder finds it.

What goes wrong on open-ended prompts

Holtzman et al.’s 2019 paper The Curious Case of Neural Text Degeneration (ICLR 2020) is the canonical write-up. They showed something weird: if you take a strong language model and ask it for the most likely continuation of a prompt, you get text that loops. Schematically (paraphrased — not a quote from the paper):

The unicorns were extremely friendly. The unicorns were extremely friendly. The unicorns were extremely friendly…

This isn’t a bug in beam search. The model genuinely assigns higher probability to the looping text than to a coherent paragraph. Why the distribution is shaped that way is still debated — Holtzman et al. document the effect; later work such as Finlayson et al. (ICLR 2024) analyses degeneration through the softmax bottleneck, though I’d stop short of saying the field has settled on a single cause. What’s solid is the observational fact: the argmax is degenerate, even though sampling from the same distribution produces human-like text. Sampling gives prose; maximizing gives mush.

Their conclusion is the load-bearing one: maximization is an inappropriate decoding objective for open-ended generation. The fix isn’t a smarter search; it’s a different objective. They proposed nucleus sampling (top-p), which is exposed as a knob across the mainstream inference stacks and chat APIs.

The “beam search curse” in translation, too

Go back to that Google Translate box. Even in machine translation — beam search’s home turf — the picture is messier than “more search is better.” Koehn and Knowles (2017, Six Challenges for Neural Machine Translation) documented the effect; the name beam search curse comes from later work (Yang, Huang and Ma, EMNLP 2018). Past modest beam widths, BLEU scores stop improving and eventually degrade. Larger beams find higher-probability sequences that are systematically too short relative to the reference; length-normalization heuristics push the sweet spot wider, but very large beams still hurt. The underlying fact is uncomfortable: even when MAP-decoding is roughly the right idea, doing it harder eventually hurts.

The shape of the lesson is the same in both worlds: the model’s probability-of-the-whole-sequence is not directly the thing you want to maximize.

Why sampling won

Sampling has three things going for it:

  1. It matches the model’s own objective. Models are trained to imitate the data distribution; drawing from that distribution gives outputs shaped like training data. Argmax-ing it produces a different beast — the mode, not a sample.
  2. It’s cheap. One forward pass per token, one beam. Beam search with width k roughly scales decode memory and bandwidth with k — and with KV-cache-bound serving, that’s a real cost.
  3. It composes with the modern toolkit. Temperature, top-p, and top-k all live inside the sampling frame. They give you a knob for diversity without abandoning the model’s distribution.

There are exceptions where beam search still earns its keep: machine translation systems shipping to production, speech recognition with a clear ground-truth target, constrained decoding where you genuinely need the highest-scoring valid output. But for “talk to me,” it’s rarely the default anymore.

Where this gets murky

Two boundaries this post won’t cross:

You started with beam search = search for the most probable sentence + the assumption that "most probable" means "best". Which half did this post kill? — not the search. The assumption. Google Translate can still keep it, because there really is roughly one right answer and the model piles its probability there. Your unicorn paragraph can’t, because “most probable” over a space of a million valid paragraphs is a different object entirely from “good.”

Check yourself

Before you go — suppose someone shows you a new model that, unlike today’s chat models, assigns its highest sequence probability to genuinely excellent prose rather than to loops. Would beam search now be the right decoder for a chatbot?

Answer

It would fix the degeneracy problem but not the whole case. Even with a perfectly behaved probability surface, beam search returns one answer — the mode — so every user asking the same question gets the same paragraph, and regenerate becomes useless. Open-ended generation wants a good answer drawn from a range of good answers, which is a different objective from “the single best one.” Plus the cost argument survives: width-k beam search multiplies your KV-cache footprint. So: better, but still probably not the default.

And one more, sharper: Koehn and Knowles found that in translation — where the maximization objective is basically right — bigger beams eventually hurt BLEU, with outputs coming out too short. What does that tell you about the model, as opposed to about the search?

Answer

It tells you the search is working and the model’s scoring is the problem. A wider beam is strictly better at finding high-probability sequences, so if quality drops as the beam widens, the sequences it’s newly finding must be high-probability-but-bad. Since sequence probability is a product of per-token probabilities, shorter sequences pay fewer multiplicative penalties — so the model’s probability surface has a systematic bias toward brevity that a weak search was accidentally hiding. That’s why length normalization exists, and it’s the same lesson as the LLM case in a milder form: don’t confuse the model’s score with the thing you actually want.

Going deeper