What is attention (in transformers)?
Every token in a sequence gets to peek at every other token and decide which ones matter. That trick is the engine inside every modern LLM.
On this page
The picture version
Attention in six pictures, for a reader who has never seen the word used this way. The prose below fills in the seams the pictures skip.
1 · The problem
One little word, two places it could point.
2 · The old way
Before attention, everything squeezed through one vector.
3 · The naive fix
Let every word see every word — but not like this.
4 · The trick
Every word holds up three different signs.
5 · The seam
Nothing is skipped. That's the whole cost.
6 · Keep this card
The whole thing on one index card.
Why it exists
Read this sentence: “The cat sat on the mat because it was warm.” Most readers land on the mat, and if you hesitated for half a second, that hesitation is the interesting part. Grammar doesn’t settle it — both nouns are available — so something else did: whatever you know about sunny floors and warm animals. And whatever that process was, it clearly involved going back to two particular earlier words rather than consulting a summary of the sentence so far. (That’s a description of the problem, not a claim about how your brain works.)
Going back to specific earlier words is exactly what machines were bad at. The standard neural approach to sequences before attention was a recurrent network: read one token, update a hidden state, read the next, update again. By the end, one fixed-size vector had to summarize everything. If the answer to “what does it refer to?” lives forty tokens back, that fact must survive forty overwrites of the same vector — and information decays. Making the network bigger doesn’t fix it; the bottleneck is the single vector everything squeezes through.
Bahdanau, Cho, and Bengio’s 2014 paper Neural Machine Translation by Jointly Learning to Align and Translate introduced the fix that stuck: let the decoder look back at every earlier state and decide, per output word, which ones matter. That mechanism is attention. Three years later, Vaswani et al.’s Attention Is All You Need (2017) asked the sharper question — if attention is doing the work, why keep the recurrence at all? — and got the transformer. We’ll build the mechanism up from that cat and that mat.
Why it matters now
Attention is the load-bearing operation inside essentially every current LLM, and inside many modern vision and multimodal models. (My read, not a citable claim: “the transformer revolution” mostly means “attention turned out to scale.”)
You meet it directly whenever you paste a long document into a chatbot. Why does the price go up faster than the length? Why is there a context limit at all, and why do providers charge less for a prompt you already sent? All three trace back to this operation: every token scores against every other token, which is on the order of N² work at length N, and the past tokens’ keys and values have to sit in GPU memory so they aren’t recomputed. KV caches, paged attention, and flash attention all exist to manage the consequences. See why attention is quadratic.
The short answer
attention = soft, content-addressed lookup over a sequence of tokens
Picture to keep: a room where every word holds up a sign saying what it is, and the word currently speaking holds up a sign saying what it’s looking for. It reads every sign at once, and copies most from whoever matches best — no searching, no order, just everyone comparing signs simultaneously.
Each token emits a query (what it’s looking for), a key (what it has), and a value (what it contributes if picked). A token’s new representation is a weighted average of everyone’s values, with weights set by how well its query matched each key. What to ask, what to advertise, and what to hand over are all learned.
How it works
Let’s try to give it some context, starting from the dumbest thing that could work, and let each version fail into the next.
Attempt 1: average everything. Replace it’s embedding with the average of all embeddings in the sentence. Now it “knows about” the cat and the mat.
Why it breaks: the average is identical for every token in the sentence. It, sat, and the all receive the same blended vector. Everything got context, and none of it is context about that token. And mat is diluted equally with the.
Attempt 2: weight the average by relevance. Instead of a flat average, weight each token by how relevant it is to the token doing the asking. Relevance is obvious enough to compute: take the dot product of the two embeddings — big when the vectors point the same way.
Why it breaks: that score is symmetric. The relevance of mat to it is forced to equal the relevance of it to mat. But those are different questions. It is looking for “a nearby noun that can be warm”; mat is advertising “I’m a flat object on a floor.” One vector per token can’t simultaneously represent what a word wants and what it offers.
Fix: split the roles. Multiply each embedding by two different learned
matrices to get a query (“what am I looking for?”) and a key (“what do
I have to offer?”). Score with Q_i · K_j. Now it’s query can point toward
warm-able objects while mat’s key advertises being one, and the score is
asymmetric — as it should be.
Why it still breaks: what do we actually average? If we average the keys, then whatever a token advertises is also everything it can contribute. A word may be easy to find by one property and useful for a completely different one.
Fix: a third projection. Add a value matrix: what the token hands over once selected, decoupled from what made it findable. Three roles, three learned matrices, per token.
But the raw scores are unbounded numbers, and we need weights that sum to 1. Pass them through a softmax.
But softmax has a failure mode. The paper’s stated reasoning is a suspicion, not a proof: they suspect that for large key dimensions the dot products grow large in magnitude, pushing softmax into regions where its gradients are extremely small — and a layer with no gradient stops learning. Their fix is a one-character one: divide the scores by √d, the square root of the key dimension, before the softmax. That’s the whole “scaled” in scaled dot-product attention:
output_i = Σ_j softmax(Q_i · K_j / √d) · V_j
The story we’re telling about our sentence is that it’s query scores high against mat’s key, so mat’s value flows into it’s new representation — and that nobody wrote a coreference rule for this; the weights that produce it were shaped only by the training objective. Treat that as an illustration of the mechanism, not a report: I’m not claiming a specific head in a specific model does this. What’s actually established is the shape — attention lets one position pull selectively from another, and the selection is learned rather than specified.
But one set of Q/K/V projections gives the model one learned way of scoring relevance at a time. Our sentence needs both “which noun does it refer to?” and “what’s the verb?” — different relations.
Fix: multiple heads. Multi-head attention runs several attention computations in parallel with different projections and concatenates the results, so different heads can carry different relations.
But now every token can see every other token — including tokens after it. A model being trained to predict the next token would just read the answer.
Fix: the causal mask. In the decoder-only models that chat assistants are built from, positions after i are masked out before the softmax when computing position i. Models with a different objective don’t need it — BERT-style encoders predict masked-out tokens rather than the next one, so seeing both directions is the point rather than a leak.
One honest note on the name. This isn’t a model of human attention; people don’t softmax over their visual field, and the going-back-to-earlier-words framing in the hook is how it feels, not what the mechanism does. Every position is compared to every other position, always, and “focus” is only the weight distribution coming out peaked — which is what the published attention heatmaps show, and plausibly why the name stuck.
You started with attention = soft, content-addressed lookup. What did this post
add that “lookup” hides? — three separate learned roles. A dictionary lookup
compares a key to a key and returns the value stored there; attention lets a
token ask with one vector, be found by another, and hand over a third. That
split is why it can learn relations nobody specified.
Famous related terms
- Self-attention —
self-attention = attention where Q, K, V all come from the same sequence— the default inside an LLM block. - Cross-attention —
cross-attention = Q from one stream + K, V from another— how a decoder reads an encoder, or a text decoder reads image features. - Multi-head attention —
MHA = N parallel attention heads, concatenated— see why MLA replaced MHA for what came next. - KV cache —
KV cache = stored K and V tensors for past tokens, reused on each decode step— turns per-step decode from O(N²) into O(N), at the cost of a lot of memory. - Quadratic cost —
attention cost ∝ N² · d— every pair of tokens gets a score; the matrix is intrinsic to the operation. - Flash attention —
FlashAttention = exact softmax attention + tiling so the N×N matrix never hits HBM— same FLOPs, far fewer memory reads.
Going deeper
- Vaswani et al., Attention Is All You Need (2017) — §3.2 is barely a page and answers “what exactly is the formula, including why the √d is there.”
- Jay Alammar, The Illustrated Transformer — read this if you want to see the matrices move; it’s the visual walkthrough most practitioners learned attention from.
- Bahdanau, Cho, Bengio (2014) — for the reader curious what attention looked like before recurrence was dropped, when it was still an add-on to an RNN translator.
A note on what I’m sure of: the Q/K/V mechanism, the role of softmax, the scaling factor’s motivation, the asymptotic cost, and the historical sequence (Bahdanau 2014 → Vaswani 2017) are all well established. Why multi-head specifically helps — what heads end up specializing in, and whether tidy “syntax head” / “coreference head” stories generalize — is more contested than a clean post like this can convey. Treat the head-specialization intuition, and my cat-and-mat walkthrough of what a head does, as a plausible sketch rather than a measured claim about any specific model.