Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why do attention sinks exist?

Trained transformers funnel a startling fraction of their attention onto the very first token — a token that's usually semantically meaningless. The pattern looks like a bug, behaves like a feature, and falls out cleanly from one constraint in the softmax.

AI & ML intermediate Apr 30, 2026 · updated Aug 25, 2026 · 12 min read

On this page

The picture version

Six pictures for a reader who knows nothing about how a chatbot handles a long conversation. The prose below fills in the seams the pictures skip.

1 · The problem

Forgetting the oldest messages is the obvious fix. It destroys the model.

“this conversation is too long” — so just drop the oldest bits the first few messages the window slides forward nonsense fine the exact moment the first messages leave the window
Sliding a window over a conversation should be harmless — the middle of a long chat drops out with no trouble. It breaks at one precise moment: when the very first tokens fall out, and then the output stops making sense.

2 · The suspect

The token it can’t live without is the one that means nothing.

a huge share token 0 a start marker every actual word of your conversation The model pours attention onto a token that carries no meaning at all — and does it in layer after layer. a beginning-of-sequence marker, or a scrap of boilerplate nothing you’d need to “look at” to predict word ten thousand
Across many layers and many heads of a trained model, a startling slice of attention — sometimes the majority — lands on the very first token. That token is usually a meaningless marker, which is what makes the pattern look like a bug.

3 · The constraint

Every part of the model must vote, every single turn. Abstaining is illegal.

one part of the model, this turn “I look for matching brackets. There are no brackets here.” it would like to sit this one out not allowed the rule it lives under all its attention must add up to exactly 1 the step that turns raw scores into attention always normalises them to a full unit so it must spend that unit somewhere There is no ballot option for “none of the above.”
The step that turns raw scores into attention weights forces them to sum to one. A head with nothing useful to look at still has to spend its full unit of attention on something — it cannot output nothing.

4 · The trick, and why evicting it is fatal

So they all dump their vote on the same harmless candidate.

token 0 is still there heads with nothing to say token 0 what it holds is close to empty, so the votes really do go nowhere a drain, not a filing cabinet token 0 has been evicted gone real words the same full unit of attention still has to go somewhere — now it floods onto text the head has no business reading and the output stops making sense
Pretraining finds a token that is always present, always boring and easy to spot, and makes it the place to pour unwanted attention. The sink is a drain, not storage — and a drain still has to exist, which is exactly why evicting it is fatal.

5 · The fix

Two ways out: pin the drain, or finally allow a blank ballot.

the mechanical fix: pin them kept forever dropped the sliding window no retraining — just a different eviction rule in the original experiments four pinned tokens were enough to keep the model coherent over very long streams the architectural fix: allow abstention none of the above real tokens no token at all add a slot in the sum that isn’t any token if a head has nothing to say, the vote goes there and its contribution really is nothing The model invented sinks because we never gave it a way to abstain.
Keeping the first few tokens pinned costs nothing and rescues a sliding window; building an explicit abstain slot into the model removes the need for a sink altogether. At least one recent open-weights model ships that slot as a trained, per-head parameter, though no public head-to-head benchmark has quantified the gain.

6 · Keep this card

The whole thing on one index card.

attention sink = a token that soaks up leftover attention + because the maths forces every head to pick somewhere it is a drain the model pours into, not storage it reads from which is why cutting the middle out of a long chat is safe, and cutting the first few tokens is not
Picture to keep: a voting booth where abstaining is illegal, so everyone with no real preference dumps their ballot on the same harmless candidate at position 0. Exactly which heads dump there, and how sinks relate to the numerical outliers that wreck low-bit quantization, is still open research.

Why it exists

You’ve probably hit the wall in a long chat session: “this conversation is too long.” The obvious engineering fix is the one you’d invent yourself — just forget the oldest messages and keep going, the way your own memory of a three-hour meeting quietly drops the first ten minutes. That is roughly what Xiao, Tian, Chen, Han and Lewis tried in their 2023 StreamingLLM work, sliding a window over an LLM’s KV cache so it could stream forever. It broke — and it broke at a very specific moment: once the first tokens of the sequence left the window, perplexity spiked and the output stopped being coherent.

Hold that scene; it’s the running example for the whole post. The first token here is usually nothing — a beginning-of-sequence marker like <s> or <|begin_of_text|>, or a scrap of system-prompt boilerplate. Nothing a model could need to “look at” to predict word ten thousand. Yet across many layers and many heads of a trained model, a huge slice of attention probability — sometimes the majority — lands right there. Xiao et al. named the pattern attention sinks in the StreamingLLM paper (ICLR 2024).

Before reading on, take a guess: what job could a semantically empty token possibly be doing that makes the model fall apart without it?

Why it matters now

Three reasons attention sinks moved from curiosity to “thing inference engineers have to know about”:

  1. Streaming and long-context generation. Any system that wants to evict old KV-cache entries to bound memory has to either preserve the sink tokens or accept that quality will collapse. StreamingLLM’s recipe — keep the first few tokens forever, slide a window over the rest — works precisely because it keeps the sinks alive.
  2. Quantization and pruning. The sink positions tend to host massive activations — numerically huge values that sit badly with quantization schemes assuming a narrow, well-behaved range. The link between attention outliers and quantization difficulty is well documented (Bondarenko et al., NeurIPS 2023); exactly how much low-bit damage is attributable to sinks specifically varies by model and I wouldn’t state a single number.
  3. Architectural fixes. At least one recent open-weights release ships the sink mechanism built into the architecture rather than letting it emerge by accident. OpenAI’s gpt-oss (released August 2025) adds a learned per-head bias logit that sits in the softmax denominator — an explicit “park your unused attention here” slot. You can see the learned per-head sink parameter directly in the reference implementation.

The phenomenon also has a neat tie to a 2023 blog post by Evan Miller, Attention Is Off By One, which proposed the same idea in spirit (softmax₁, with an extra +1 in the denominator) on theoretical grounds, before the StreamingLLM paper showed how badly real models need somewhere to abstain.

The short answer

attention sink = a token that absorbs leftover attention weight + because softmax forces every head to pick somewhere

Picture to keep: a voting booth where abstaining is illegal — every head must cast its full ballot each turn, so the ones with no real preference all dump their vote on the same harmless candidate sitting at position 0.

Softmax outputs always sum to 1. So every attention head, on every token, on every layer, must spend its full unit of attention on something — even when it has nothing useful to attend to. Trained models discover that the cheapest place to dump that surplus is a position that’s reliably present in every sequence and reliably uninformative: the first token. Sinks are the model’s “no-op” hack, forced into existence by the sum-to-one constraint.

How it works

Start from the attention formula. For a query $q_t$ at position $t$ and keys $k_0, \dots, k_t$:

weights = softmax(q_t · K^T / sqrt(d))
output  = weights · V

Softmax means weights[i] = exp(score[i]) / Σ_j exp(score[j]). The weights are non-negative and sum to exactly 1.

Now imagine you’re a particular attention head, and on this particular token, you genuinely have nothing to contribute. Maybe you’re a head that specializes in “find the matching open-paren” and there are no parens around. You’d like to abstain — output a zero vector, contribute nothing to the residual stream — but softmax will not let you. You have to put your full unit of probability mass somewhere. Whichever key you score highest wins, even if all your scores are tiny.

The standard account of what happens next — and I’d call it the dominant interpretation rather than a proven mechanism — is that pretraining finds a stable trick: pick a token that is (a) reliably present in every sequence, (b) reliably boring, and (c) easy to identify. The first token fits all three. Concentrate attention there when there’s nothing else to say. If the value vector at that position carries little information, the model effectively gets its “abstain” back. This is the sink. The voting-booth analogy breaks here in a useful way: a real harmless candidate would still take office with all those votes, whereas the sink’s value contributes close to nothing downstream, so the votes really do vanish. That’s the whole trick — and it’s also why sliding the sink out of the window is fatal. Evict it and every abstaining head is forced to redistribute its full unit of attention onto real tokens it has no business reading.

A few details worth showing the seams on:

The “off by one” framing

Evan Miller’s 2023 post made the same point from the algebra side without naming attention sinks. His proposal — softmax₁(x)_i = exp(x_i) / (1 + Σ_j exp(x_j)) — adds a phantom 1 to the denominator. The phantom term effectively gives every head a free “attend to nothing” option that doesn’t correspond to any real token. If all real scores are very negative, the phantom dominates, the real weights all go to nearly zero, and the head genuinely abstains. The output for that head is then close to the zero vector.

gpt-oss ships a per-head learnable version of this idea: each head has its own bias logit appended to the attention scores, and that logit is trained jointly with the rest of the model. If a head wants to “park” a lot of attention on the bias slot, it learns a high value; if it always wants to attend to real tokens, it learns a low one. This is, in effect, the architectural answer to “why did the model invent attention sinks?”: because we never gave it a legitimate way to abstain. Now we do.

There is no clean apples-to-apples public benchmark establishing how much the learned-bias approach beats a comparable model trained without it — the ablation would require training two models at scale, and nobody has published one. The intuition is strong; the quantified case is still being assembled.

What still puzzles people

A few honest gaps:

You started with attention sink = a token that absorbs leftover attention weight. What did this post add about why that token can be evicted safely — or not? — + because softmax forces every head to pick somewhere. The sink isn’t storage the model is reading from; it’s a drain the model is pouring into, and a drain still has to exist. That’s why sliding a window over your chat history is safe for the middle and fatal for the first four tokens.

Check yourself

Before you go — suppose you’re serving a model with a sliding-window KV cache and you pin the first 4 tokens, as StreamingLLM prescribes. A colleague suggests saving memory by pinning only the keys for those tokens and dropping their value vectors, since the post said the sink’s values are near-uninformative. Would you sign off?

Answer

No — or at least not on that reasoning. “Near-uninformative” is a claim about what the trained model drove the values toward, not a guarantee they’re exactly zero, and it’s model-specific. More importantly, the attention output is a weighted sum of value vectors; if you drop the sink’s value you have to decide what to substitute, and any nonzero substitute gets multiplied by a very large weight (that’s the whole point of a sink) and injected into the residual stream. The saving is also tiny: four tokens out of a window of thousands. This is the kind of “obviously equivalent” optimization that needs a perplexity measurement, not an argument.

And one more: a model is pretrained with a guaranteed <s> token at the start of every document. Would you expect it to need more sink tokens than a model trained on documents with inconsistent openings, or fewer?

Answer

Fewer — that’s the Xiao et al. conjecture. If position 0 is reliably the same token in every training document, the model can learn one unambiguous “dump here” address. Without that consistency it has to treat several early positions as usable sinks, which is why keeping just one initial token wasn’t enough in their experiments and four was. Note this is a conjecture with supporting experiments, not a proven law — you’d want to measure it per model family rather than assume it.

Going deeper