Why do attention sinks exist?
Trained transformers funnel a startling fraction of their attention onto the very first token — a token that's usually semantically meaningless. The pattern looks like a bug, behaves like a feature, and falls out cleanly from one constraint in the softmax.
On this page
The picture version
Six pictures for a reader who knows nothing about how a chatbot handles a long conversation. The prose below fills in the seams the pictures skip.
1 · The problem
Forgetting the oldest messages is the obvious fix. It destroys the model.
2 · The suspect
The token it can’t live without is the one that means nothing.
3 · The constraint
Every part of the model must vote, every single turn. Abstaining is illegal.
4 · The trick, and why evicting it is fatal
So they all dump their vote on the same harmless candidate.
5 · The fix
Two ways out: pin the drain, or finally allow a blank ballot.
6 · Keep this card
The whole thing on one index card.
Why it exists
You’ve probably hit the wall in a long chat session: “this conversation is too long.” The obvious engineering fix is the one you’d invent yourself — just forget the oldest messages and keep going, the way your own memory of a three-hour meeting quietly drops the first ten minutes. That is roughly what Xiao, Tian, Chen, Han and Lewis tried in their 2023 StreamingLLM work, sliding a window over an LLM’s KV cache so it could stream forever. It broke — and it broke at a very specific moment: once the first tokens of the sequence left the window, perplexity spiked and the output stopped being coherent.
Hold that scene; it’s the running example for the whole post. The first
token here is usually nothing — a beginning-of-sequence marker like
<s> or <|begin_of_text|>, or a scrap of system-prompt boilerplate.
Nothing a model could need to “look at” to predict word ten thousand.
Yet across many layers and many heads of a trained model, a huge slice
of attention probability — sometimes the majority — lands right there.
Xiao et al. named the pattern attention sinks in the StreamingLLM
paper (ICLR 2024).
Before reading on, take a guess: what job could a semantically empty token possibly be doing that makes the model fall apart without it?
Why it matters now
Three reasons attention sinks moved from curiosity to “thing inference engineers have to know about”:
- Streaming and long-context generation. Any system that wants to evict old KV-cache entries to bound memory has to either preserve the sink tokens or accept that quality will collapse. StreamingLLM’s recipe — keep the first few tokens forever, slide a window over the rest — works precisely because it keeps the sinks alive.
- Quantization and pruning. The sink positions tend to host massive activations — numerically huge values that sit badly with quantization schemes assuming a narrow, well-behaved range. The link between attention outliers and quantization difficulty is well documented (Bondarenko et al., NeurIPS 2023); exactly how much low-bit damage is attributable to sinks specifically varies by model and I wouldn’t state a single number.
- Architectural fixes. At least one recent open-weights release
ships the sink mechanism built into the architecture rather than
letting it emerge by accident. OpenAI’s
gpt-oss(released August 2025) adds a learned per-head bias logit that sits in the softmax denominator — an explicit “park your unused attention here” slot. You can see the learned per-head sink parameter directly in the reference implementation.
The phenomenon also has a neat tie to a 2023 blog post by Evan Miller,
Attention Is Off By One, which proposed the same idea in spirit
(softmax₁, with an extra +1 in the denominator) on theoretical
grounds, before the StreamingLLM paper showed how badly real models
need somewhere to abstain.
The short answer
attention sink = a token that absorbs leftover attention weight + because softmax forces every head to pick somewhere
Picture to keep: a voting booth where abstaining is illegal — every head must cast its full ballot each turn, so the ones with no real preference all dump their vote on the same harmless candidate sitting at position 0.
Softmax outputs always sum to 1. So every attention head, on every token, on every layer, must spend its full unit of attention on something — even when it has nothing useful to attend to. Trained models discover that the cheapest place to dump that surplus is a position that’s reliably present in every sequence and reliably uninformative: the first token. Sinks are the model’s “no-op” hack, forced into existence by the sum-to-one constraint.
How it works
Start from the attention formula. For a query $q_t$ at position $t$ and keys $k_0, \dots, k_t$:
weights = softmax(q_t · K^T / sqrt(d))
output = weights · V
Softmax means weights[i] = exp(score[i]) / Σ_j exp(score[j]). The
weights are non-negative and sum to exactly 1.
Now imagine you’re a particular attention head, and on this particular token, you genuinely have nothing to contribute. Maybe you’re a head that specializes in “find the matching open-paren” and there are no parens around. You’d like to abstain — output a zero vector, contribute nothing to the residual stream — but softmax will not let you. You have to put your full unit of probability mass somewhere. Whichever key you score highest wins, even if all your scores are tiny.
The standard account of what happens next — and I’d call it the dominant interpretation rather than a proven mechanism — is that pretraining finds a stable trick: pick a token that is (a) reliably present in every sequence, (b) reliably boring, and (c) easy to identify. The first token fits all three. Concentrate attention there when there’s nothing else to say. If the value vector at that position carries little information, the model effectively gets its “abstain” back. This is the sink. The voting-booth analogy breaks here in a useful way: a real harmless candidate would still take office with all those votes, whereas the sink’s value contributes close to nothing downstream, so the votes really do vanish. That’s the whole trick — and it’s also why sliding the sink out of the window is fatal. Evict it and every abstaining head is forced to redistribute its full unit of attention onto real tokens it has no business reading.
A few details worth showing the seams on:
- It’s not always position 0. The Xiao et al. paper tested keeping just one initial token and found one wasn’t enough — four worked. The reason they conjecture: the models they studied weren’t pretrained with a consistent first token across all training documents, so the model learned to use several early positions as sinks rather than just one. Models pretrained with a guaranteed start-of-sequence token may concentrate on a single sink instead. This is empirical; the exact number is model-specific.
- Sinks correlate with massive activations. This next part is interpretation, not settled fact. Subsequent work observed that the hidden-state norms at sink positions can be orders of magnitude larger than at other positions (Sun et al. 2024, Massive Activations in Large Language Models). The common reading is that these large activations are how the model encodes “this is the sink” robustly enough that all the heads can find it — plausible, and the mechanistic picture is still being filled in.
- The fix is mechanical. StreamingLLM doesn’t retrain anything. It just changes the KV-cache eviction policy: keep the first k tokens pinned forever (their experiments settled on 4), and slide a window over everything else. With the sinks preserved they report stable language modelling out to roughly 4 million tokens; without them, the naive sliding window degrades far sooner. Go to the paper’s figures for the exact curves rather than trusting a summary sentence.
The “off by one” framing
Evan Miller’s 2023 post made the same point from the algebra side
without naming attention sinks. His proposal — softmax₁(x)_i = exp(x_i) / (1 + Σ_j exp(x_j)) —
adds a phantom 1 to the denominator. The phantom term effectively
gives every head a free “attend to nothing” option that doesn’t
correspond to any real token. If all real scores are very negative,
the phantom dominates, the real weights all go to nearly zero, and the
head genuinely abstains. The output for that head is then close to
the zero vector.
gpt-oss ships a per-head learnable version of this idea: each head
has its own bias logit appended to the attention scores, and that
logit is trained jointly with the rest of the model. If a head wants
to “park” a lot of attention on the bias slot, it learns a high
value; if it always wants to attend to real tokens, it learns a low
one. This is, in effect, the architectural answer to “why did the
model invent attention sinks?”: because we never gave it a legitimate
way to abstain. Now we do.
There is no clean apples-to-apples public benchmark establishing how much the learned-bias approach beats a comparable model trained without it — the ablation would require training two models at scale, and nobody has published one. The intuition is strong; the quantified case is still being assembled.
What still puzzles people
A few honest gaps:
- Why does the model need several sink tokens rather than one, even when sequences begin with a consistent BOS marker? Xiao et al. conjecture it’s pretraining-data dependent; that hasn’t been pinned down rigorously across model families.
- Sinks behave differently across heads, layers, and positions. There isn’t yet a unified mechanistic story that explains which heads dump to the sink and when. This is still open research; the Barbero et al. paper below is one recent attempt.
- The connection between attention sinks and outlier features in quantization is empirically suggestive but not fully explained, and the causal direction is contested — do sinks cause the outliers that wreck low-bit quantization, or are both downstream of something else? No published result settles it.
You started with attention sink = a token that absorbs leftover attention weight. What did this post add about why that token can be
evicted safely — or not? — + because softmax forces every head to pick somewhere. The sink isn’t storage the model is reading from; it’s a
drain the model is pouring into, and a drain still has to exist. That’s
why sliding a window over your chat history is safe for the middle and
fatal for the first four tokens.
Check yourself
Before you go — suppose you’re serving a model with a sliding-window KV cache and you pin the first 4 tokens, as StreamingLLM prescribes. A colleague suggests saving memory by pinning only the keys for those tokens and dropping their value vectors, since the post said the sink’s values are near-uninformative. Would you sign off?
Answer
No — or at least not on that reasoning. “Near-uninformative” is a claim about what the trained model drove the values toward, not a guarantee they’re exactly zero, and it’s model-specific. More importantly, the attention output is a weighted sum of value vectors; if you drop the sink’s value you have to decide what to substitute, and any nonzero substitute gets multiplied by a very large weight (that’s the whole point of a sink) and injected into the residual stream. The saving is also tiny: four tokens out of a window of thousands. This is the kind of “obviously equivalent” optimization that needs a perplexity measurement, not an argument.
And one more: a model is pretrained with a guaranteed <s> token at
the start of every document. Would you expect it to need more sink
tokens than a model trained on documents with inconsistent openings, or
fewer?
Answer
Fewer — that’s the Xiao et al. conjecture. If position 0 is reliably the same token in every training document, the model can learn one unambiguous “dump here” address. Without that consistency it has to treat several early positions as usable sinks, which is why keeping just one initial token wasn’t enough in their experiments and four was. Note this is a conjecture with supporting experiments, not a proven law — you’d want to measure it per model family rather than assume it.
Famous related terms
- Softmax —
softmax(x)_i = exp(x_i) / Σ_j exp(x_j)— the sum-to-one constraint that creates the problem in the first place. softmax₁/ “Attention Is Off By One” —softmax₁ = softmax + phantom 1 in the denominator— Evan Miller’s proposed fix; lets a head abstain.- StreamingLLM —
StreamingLLM = sliding-window KV cache + pinned first-k tokens— Xiao et al.’s recipe for unbounded-length generation without retraining. - Massive activations —
massive activations = hidden states with anomalously huge norms— empirically co-located with sink positions; not yet fully explained. - KV cache —
KV cache = stored attention keys and values + reused on every decode step— the structure whose eviction policy made the sink visible. - PagedAttention —
PagedAttention = block-based KV cache + page table— orthogonal to sinks, but real systems have to combine the two: pin the sink blocks, page everything else.
Going deeper
- Xiao, Tian, Chen, Han, Lewis — Efficient Streaming Language Models with Attention Sinks (ICLR 2024), with code — start here if you want the actual measurements behind “four tokens, pinned forever.”
- Evan Miller — Attention Is Off By One (July 2023) — the clearest walk-through of why the softmax denominator is the culprit, argued from algebra before anyone had the streaming evidence.
- Barbero, Arroyo, Gu, Perivolaropoulos, Bronstein, Veličković, Pascanu — Why do LLMs attend to the first token? (2025) — the rabbit hole, for readers who want the current attempt at a mechanistic account rather than a conjecture.