Why beam search died for LLMs
Beam search was the default way to decode neural sequence models for years. Then chatbots arrived and quietly stopped using it. The reason is stranger than 'sampling is more creative.'
On this page
The picture version
Six pictures for a reader who has never chosen a decoding strategy. The prose below fills in the seams the pictures skip.
1 · The problem
One box repeats itself on purpose. The other refuses to.
2 · What beam search is
Keep the best few half-sentences alive instead of committing to one.
3 · The surprise
Search harder and the writing gets worse.
4 · Why
The peak of the landscape is not where the good writing lives.
5 · Even on home turf
Translation has the same disease, in a milder form.
6 · Keep this card
Which half of the definition died.
Why it exists
Paste a sentence into Google Translate, hit translate twice, and you get the same output both times. Ask a chatbot to write you a paragraph about unicorns, hit regenerate, and you get a different paragraph. Same underlying kind of model — a neural net predicting one token at a time — and yet one behaves like a lookup and the other behaves like a die roll. That difference isn’t an accident of branding. It’s a decision about how to walk the model’s probabilities, and the two systems make opposite choices. Keep both cases in mind; the whole post is about why the second one had to abandon the first one’s method.
The first one’s method is beam search. Through the mid-2010s it was the workhorse decoder for NMT, summarization, and image captioning. The intuition is clean: greedy decoding picks the locally best token and gets stuck; beam search keeps the top-k partial sequences alive at every step, so it has a shot at finding a globally higher-probability sequence. More search, better answer.
Then LLMs arrived, ChatGPT shipped, and beam search quietly stopped being the default for open-ended generation. The OpenAI chat completions API doesn’t expose it; what it exposes are sampling knobs — temperature and top-p — and no beam width.
That’s the puzzle. You probably assume beam search fell out of favour because sampling is “more creative” — a deliberate trade of quality for variety. It’s very nearly the opposite: on open-ended prompts, the harder you search for the model’s best answer, the worse the text you get.
Why it matters now
Every time you call a chat model, something is choosing one of tens of thousands of tokens per step. That choice is a strategy, not a fact about the model. Picking the wrong strategy makes a state-of-the-art model produce slop — bland paragraphs, repetitive loops, weirdly short answers. Engineers who reach for beam search expecting “higher quality output” get the opposite, and the failure mode looks like the model itself is bad.
The shift also reshaped how people think about prompting. If decoding is sampling, then “the answer” isn’t a thing the model has — it’s a distribution the model has, and you’re drawing from it. That mental model is load-bearing for understanding why the same prompt gives different outputs, why temperature matters, and why two perfectly correct answers can sit side by side.
The short answer
beam search = search for the most probable sentence + the assumption that "most probable" means "best"
Picture to keep: a crowd of a million people who each wrote one paragraph about unicorns. Sampling hands you a random paragraph from the pile. Beam search hunts for the single paragraph the crowd was most likely to have produced — and that turns out to be one dull sentence copy-pasted six times. The analogy breaks in one place worth naming: the model isn’t averaging anyone’s paragraphs. It’s assigning a probability to every possible paragraph, and the peak of that landscape simply isn’t where the good writing lives.
Beam search optimizes for the highest-probability sequence under the model. That assumption holds when there is one right answer and the model concentrates its probability near it — translating your sentence, transcribing speech. It fails when there are many valid continuations, because then the highest-probability sequence is a degenerate one nobody wanted. For open-ended generation, the most probable output is worse than a randomly sampled one. So sampling won.
How it works
The mechanism is counter-intuitive enough that it’s worth walking through.
Beam search, briefly
At each step, keep the top-k partial sequences (the “beam”) ranked by cumulative log-probability. Expand each by one token, score all kN candidates, keep the top k again. At the end, return the highest-scoring full sequence. With k=1 you get greedy decoding; as k grows, you approach exact MAP decoding (true exhaustive argmax over sequences requires more than just a wide beam, but the intuition is right).
For a translation system, this is great. There’s roughly one correct translation; the model concentrates probability on tokens near that translation; searching harder finds it.
What goes wrong on open-ended prompts
Holtzman et al.’s 2019 paper The Curious Case of Neural Text Degeneration (ICLR 2020) is the canonical write-up. They showed something weird: if you take a strong language model and ask it for the most likely continuation of a prompt, you get text that loops. Schematically (paraphrased — not a quote from the paper):
The unicorns were extremely friendly. The unicorns were extremely friendly. The unicorns were extremely friendly…
This isn’t a bug in beam search. The model genuinely assigns higher probability to the looping text than to a coherent paragraph. Why the distribution is shaped that way is still debated — Holtzman et al. document the effect; later work such as Finlayson et al. (ICLR 2024) analyses degeneration through the softmax bottleneck, though I’d stop short of saying the field has settled on a single cause. What’s solid is the observational fact: the argmax is degenerate, even though sampling from the same distribution produces human-like text. Sampling gives prose; maximizing gives mush.
Their conclusion is the load-bearing one: maximization is an inappropriate decoding objective for open-ended generation. The fix isn’t a smarter search; it’s a different objective. They proposed nucleus sampling (top-p), which is exposed as a knob across the mainstream inference stacks and chat APIs.
The “beam search curse” in translation, too
Go back to that Google Translate box. Even in machine translation — beam search’s home turf — the picture is messier than “more search is better.” Koehn and Knowles (2017, Six Challenges for Neural Machine Translation) documented the effect; the name beam search curse comes from later work (Yang, Huang and Ma, EMNLP 2018). Past modest beam widths, BLEU scores stop improving and eventually degrade. Larger beams find higher-probability sequences that are systematically too short relative to the reference; length-normalization heuristics push the sweet spot wider, but very large beams still hurt. The underlying fact is uncomfortable: even when MAP-decoding is roughly the right idea, doing it harder eventually hurts.
The shape of the lesson is the same in both worlds: the model’s probability-of-the-whole-sequence is not directly the thing you want to maximize.
Why sampling won
Sampling has three things going for it:
- It matches the model’s own objective. Models are trained to imitate the data distribution; drawing from that distribution gives outputs shaped like training data. Argmax-ing it produces a different beast — the mode, not a sample.
- It’s cheap. One forward pass per token, one beam. Beam search with width k roughly scales decode memory and bandwidth with k — and with KV-cache-bound serving, that’s a real cost.
- It composes with the modern toolkit. Temperature, top-p, and top-k all live inside the sampling frame. They give you a knob for diversity without abandoning the model’s distribution.
There are exceptions where beam search still earns its keep: machine translation systems shipping to production, speech recognition with a clear ground-truth target, constrained decoding where you genuinely need the highest-scoring valid output. But for “talk to me,” it’s rarely the default anymore.
Where this gets murky
Two boundaries this post won’t cross:
- The cause is still open. Holtzman et al. document that the argmax is degenerate; they don’t settle why. Finlayson et al.’s Closing the Curious Case of Neural Text Degeneration (ICLR 2024) argues the softmax bottleneck is the culprit, but the field hasn’t converged on one explanation, and this post takes the effect as the solid part and the mechanism as contested.
- “Not the default” is narrower than “gone.” The mainstream chat APIs expose temperature and top-p and no beam width. Open-source inference engines (Hugging Face Transformers, vLLM, and others) still ship beam-search implementations — they’re just not what people reach for when serving a chatbot.
You started with beam search = search for the most probable sentence + the assumption that "most probable" means "best". Which half did this post kill? — not the search. The assumption. Google Translate can still keep it, because there really is roughly one right answer and the model piles its probability there. Your unicorn paragraph can’t, because “most probable” over a space of a million valid paragraphs is a different object entirely from “good.”
Check yourself
Before you go — suppose someone shows you a new model that, unlike today’s chat models, assigns its highest sequence probability to genuinely excellent prose rather than to loops. Would beam search now be the right decoder for a chatbot?
Answer
It would fix the degeneracy problem but not the whole case. Even with a perfectly behaved probability surface, beam search returns one answer — the mode — so every user asking the same question gets the same paragraph, and regenerate becomes useless. Open-ended generation wants a good answer drawn from a range of good answers, which is a different objective from “the single best one.” Plus the cost argument survives: width-k beam search multiplies your KV-cache footprint. So: better, but still probably not the default.
And one more, sharper: Koehn and Knowles found that in translation — where the maximization objective is basically right — bigger beams eventually hurt BLEU, with outputs coming out too short. What does that tell you about the model, as opposed to about the search?
Answer
It tells you the search is working and the model’s scoring is the problem. A wider beam is strictly better at finding high-probability sequences, so if quality drops as the beam widens, the sequences it’s newly finding must be high-probability-but-bad. Since sequence probability is a product of per-token probabilities, shorter sequences pay fewer multiplicative penalties — so the model’s probability surface has a systematic bias toward brevity that a weak search was accidentally hiding. That’s why length normalization exists, and it’s the same lesson as the LLM case in a milder form: don’t confuse the model’s score with the thing you actually want.
Famous related terms
- Greedy decoding —
greedy = beam search + k=1— pick the argmax token each step. Fast, deterministic, prone to repetition. - Nucleus sampling (top-p) —
top-p = sampling + truncate the tail to mass p— the Holtzman et al. fix; default in most modern stacks. - Top-k sampling —
top-k = sampling + keep only k highest-probability tokens— older cousin of top-p, less adaptive. - Temperature —
temperature = scale logits before softmax— the diversity dial that lives inside the sampling frame. - MAP decoding —
MAP = argmax over the whole sequence— what beam search is approximating, and what turns out to be the wrong objective for open-ended text. - Minimum Bayes Risk (MBR) decoding —
MBR = pick the candidate that minimizes expected loss vs. the model's distribution— in practice approximated by sampling N candidates and returning the one most similar (by BLEU/COMET) to the others. Seeing renewed interest as a sampling-era alternative to beam search in MT.
Going deeper
- Holtzman, Buys, Du, Forbes, Choi (2019), The Curious Case of Neural Text Degeneration — the primary source, and where to go if your question is “how do they actually show that the most probable text is worse than a sample?”
- Hugging Face, How to generate text — the explainer, for “what do greedy, beam, top-k and top-p literally produce on the same prompt?” Side-by-side outputs make the difference visceral in a way prose can’t.
- Koehn and Knowles (2017), Six Challenges for Neural Machine Translation — the rabbit hole, if you want to know why even translation, beam search’s home turf, has a beam-width sweet spot.