Why do positional encodings exist?
A transformer cannot tell 'dog bites man' from 'man bites dog' on its own. The attention math is symmetric in token order — until you bolt on a position signal. Every modern LLM does, and the choice of how shapes long-context behavior more than people realize.
On this page
- The picture version
- Why it exists
- Why it matters now
- The short answer
- How it works
- What “permutation equivariant” actually buys you
- Option 1: Add position to the input (sinusoidal, learned absolute)
- Option 2: Rotate Q and K (RoPE)
- Option 3: Bias the attention scores directly (ALiBi)
- Option 4: No positional encoding at all (NoPE)
- Why long-context is mostly a RoPE-extension story
- The seams worth seeing
- Check yourself
- Famous related terms
- Going deeper
The picture version
Seven pictures for a reader who has never thought about word order as a problem. The prose below fills in the seams the pictures skip.
1 · The problem
Same six words. Attention cannot tell them apart.
2 · The first fix
Give every token a numbered jersey before it walks in.
3 · The move that won
Don’t label the token. Rotate it.
4 · Two roads not taken
Bias the scores instead — or encode nothing at all.
5 · Why long context breaks
The angles wander into phase space the model never trained on.
6 · The knobs
Every long-context recipe is a knob on the rotation schedule.
7 · Keep this card
The whole thing on one index card.
Why it exists
Type “the dog bit the man” into any chatbot, then type “the man bit the dog”, and you get two different answers — obviously, because they mean different things. Nothing about that feels like an achievement. But those two prompts are the same six words. The only thing separating them is order, and the core operation inside a transformer is mathematically blind to order. Somebody had to add the ability to tell those two sentences apart.
That’s the embarrassing fact the original transformer paper had to fix in a footnote-shaped way: attention has no idea what order its input tokens came in.
Take that sentence, shuffle the tokens, and feed both versions through pure self-attention: the math comes out the same up to a permutation of the outputs. Attention is a weighted sum over a set of (key, value) pairs. Sets have no order. “the dog bit the man” and “the man bit the dog” produce the same set, so the same attention scores, so the same internal representations — just rearranged. That property has a name: permutation equivariant. It’s exactly what you don’t want for language, where order is most of the meaning.
The original Attention Is All You Need paper (Vaswani et al., 2017) handled this by adding a fixed sinusoidal vector to each token’s input embedding — one vector per position, the same one every time, baked from sines and cosines at geometrically-spaced frequencies. Token embedding plus position vector goes into the model. That’s the entire trick.
That early choice has aged into a much bigger story. The open-weights LLM families whose architectures are public — LLaMA, Mistral, Gemma, Qwen, GPT-NeoX — don’t use the original sinusoidal scheme. They use RoPE, which works by rotating the query and key vectors inside the attention layer rather than adding anything to the embeddings. (Closed frontier models mostly don’t document their positional scheme, so treat “everyone uses RoPE” as an inference from the open ecosystem, not a surveyed fact.) The reason matters, and it’s why “long context” was hard in a way nobody quite expected.
Why it matters now
Three things make positional encoding a load-bearing choice today, not a footnote:
- Long context is the headline feature. When you advertise 128k or 1M token windows, the positions in those windows have to mean something the model was actually trained to handle. Feed a model positions far past anything it trained on and quality degrades — that’s the premise the whole context-extension literature is built on (YaRN opens with it). The game of “context extension” — YaRN, NTK scaling, position interpolation — is fundamentally about rescaling RoPE to work at lengths it never saw.
- The KV cache is positional. Each entry in the KV cache (post) is the K and V tensor for a specific position in the prompt. With RoPE, the rotation has already been applied to the cached K when it was computed. That means the position is baked into the cache; you can’t trivially “shift” a cached prefix to a new offset. This shows up in subtle ways when systems try to share or splice KV caches across requests.
- It’s the easiest place for “I’m using a long context model” to silently mean “I’m using a confused model.” Models extended past their training length without proper RoPE rescaling don’t crash — they just degrade in ways that are hard to spot from a single completion. (Lost in the middle-style pathologies are plausibly partly downstream of this, though they have other causes too and no published work pins the relative share.)
The choice of positional encoding is one of those design decisions where the wrong answer doesn’t fail loudly. It just makes the model quietly worse at long-range work.
The short answer
positional encoding = a position-dependent signal injected into Q/K (or input embeddings) so attention can tell tokens apart by where they are
Picture to keep: every token walks into the room wearing a numbered jersey — and in the modern version the number isn’t printed on the jersey at all, it’s the angle each player is turned to face, so what attention reads off is how far apart two players are pointing. Where the analogy breaks: a jersey number is one label per player, while the rotation is applied separately to dozens of 2D slices of each vector at different speeds, so there’s no single “position” written down anywhere.
Self-attention by itself sees a bag of tokens, not a sequence. To get sequence behavior you have to inject “you are token #i” somewhere in the pipeline before the attention scores get computed — into the input embeddings, or into the Q and K vectors that attention actually multiplies. The original transformer added a fixed sinusoid to each input embedding. Modern open-weights LLMs (LLaMA, Mistral, GPT-NeoX, Gemma, Qwen) instead rotate the query and key vectors inside attention — that’s RoPE — because it makes the positional part of the score depend only on the relative distance between two tokens, which generalizes and composes much better.
How it works
Start with the failure mode, because the rest follows from it.
What “permutation equivariant” actually buys you
A single attention head computes, for each token i:
output_i = Σ_j softmax_j( q_i · k_j / √d ) · v_j
Notice what this depends on: the content of q_i, and the contents of all k_j and v_j. There is nothing here about where token j sits in the sequence. If you shuffle the tokens, you shuffle the (k, v) pairs, but the set of pairs is unchanged — so for any given query, the scores it computes are unchanged (modulo which output slot you read out).
For a vision model on patches that’s sometimes fine; spatial position can be encoded in other ways. For language, it’s catastrophic: “the dog bit the man”, “the man bit the dog”, and “bit the the man dog” are all the same multiset of tokens.
You have to break the symmetry. The question is where you break it, and how.
Option 1: Add position to the input (sinusoidal, learned absolute)
The original transformer’s move: precompute a vector PE_i for each position i, and add it to the token embedding before layer 1.
x_i = embedding(token_i) + PE_i
Vaswani et al. defined PE_i using sines and cosines:
PE_{i, 2k} = sin(i / 10000^{2k/d})
PE_{i, 2k+1} = cos(i / 10000^{2k/d})
The 10000 is just a chosen base; the shape is geometric: low dimensions oscillate fast, high dimensions oscillate slowly. The reason for sines and cosines specifically (rather than learning the position vectors) was, in their own words, that it “may allow the model to extrapolate to sequence lengths longer than the ones encountered during training.” A nice hope, and one the later literature didn’t bear out: when Press et al. measured extrapolation directly for the ALiBi paper, sinusoidal models degraded past their training length rather than gracefully extending. The construction stuck for years anyway.
BERT and many follow-ups used learned absolute position embeddings instead: a lookup table from position index to a learned vector, just like the token embedding table. Same shape, simpler, no extrapolation property at all (positions past training length are literally untrained vectors).
The intuitive worry with both flavors of “add to input”: the position information enters once, at the bottom, and has to survive through every layer’s mixing. By layer 24 the model has been shuffling these representations around for a long time. Whether the signal is meaningfully attenuated in practice isn’t settled in the literature — treat it as motivation for what comes next, not as established fact. What is established is that absolute schemes encode absolute position, when what attention usually wants is relative position — token i caring about “the token three to my left” rather than “the token at index 47.”
Option 2: Rotate Q and K (RoPE)
RoPE — proposed in Su et al.’s RoFormer paper (Su, Lu, Pan, Murtadha, Wen, Liu, 2021; arXiv 2104.09864) — does something cleaner. Instead of adding to the embedding, it rotates the query and key vectors themselves, by an amount that depends on the position.
The key fact, written compactly: split each head’s d-dimensional Q and K into d/2 pairs of components. For each pair, treat it as a 2D vector and rotate it by an angle m · θ_k, where m is the position and θ_k is a per-pair frequency (using the same 10000-base geometric progression as the original sinusoidal scheme). Different pairs rotate at different rates. The whole operation is a 2×2 rotation applied independently to each pair.
The magic: the dot product of a rotated query at position m with a rotated key at position n depends on the content of q and k plus the relative offset m − n, not on the absolute m and n separately. This is a mathematical fact about rotations: R(α) · R(β)ᵀ = R(α − β). Concretely: “bit” attending to “dog” one token to its left produces the same positional contribution whether that phrase starts at token 1 or token 40,000 — and it’s still a different contribution from “bit” attending to “dog” one token to its right, which is the whole difference between our two sentences. Translate the sequence and the attention pattern stays identical; flip the order and it doesn’t.
Other useful properties of RoPE:
- No new parameters. It’s a fixed transform on Q and K. Nothing is learned that wasn’t learned before.
- Position is reinjected at every layer. Because RoPE is applied inside the attention block — every layer — there’s no concern about a single bottom-of-stack injection getting lost in the deep layers, the way there might be with additive input embeddings. (Whether that actually matters in practice is a separate empirical question, and not one the literature has cleanly answered.)
- It composes with the KV cache. The rotated K is what gets cached. Subsequent queries are rotated and dotted against the already-rotated cached keys; the relative-position math still works out.
- Decaying inter-token dependency with distance. Su et al. show that, in expectation, the attention score from RoPE has a built-in mild decay as distance increases. This isn’t a hard cutoff, but it’s a useful inductive bias.
LLaMA, Mistral, Gemma, GPT-NeoX, GPT-J, and Qwen all use RoPE — it’s the default across the open-weights families whose configs you can read. What the closed frontier labs use is mostly undocumented, so this is a claim about the open ecosystem rather than a surveyed share.
Option 3: Bias the attention scores directly (ALiBi)
A third route: don’t touch Q or K at all. Just bias the attention score for token i attending to token j by −|i − j| · m_h, where m_h is a small per-head slope. Closer tokens get a smaller penalty; farther tokens get a larger one. This is ALiBi, from Press, Smith, and Lewis (arXiv 2108.12409, ICLR 2022).
ALiBi’s pitch in the paper was specifically about extrapolation: train at length 1024, test at 2048+, with no fine-tuning. The headline result was a 1.3B model trained on 1024 tokens that extrapolated to 2048 with the same perplexity as a sinusoidal model trained at 2048, while training 11% faster and using 11% less memory. ALiBi has been used in some real models (e.g. MPT, BLOOM), but the open-weights families that came after mostly chose RoPE. One plausible read: RoPE’s relative-position-aware Q/K rotations look more like a general mechanism, where ALiBi is a hand-designed bias term that happens to work for distance-decay specifically. No public source attributes the convergence to a single reason, so treat that as a take rather than a finding.
Option 4: No positional encoding at all (NoPE)
This one is the most surprising. Haviv et al. (arXiv 2203.16634, 2022) showed that causal (decoder-only) transformers with no positional encoding at all are still competitive with explicit-PE ones at language modeling, across multiple datasets and model sizes. The intuition the paper develops is that the causal mask itself — token i can only attend to tokens 1…i — leaks position information: a token at position 5 has 5 things to attend to; a token at position 50 has 50. (Their probing experiments suggest the model picks up an implicit notion of absolute position; the exact mechanism is more subtle than a literal “count.”)
Kazemnejad et al. (The Impact of Positional Encoding on Length Generalization in Transformers, arXiv 2305.19466, 2023) studied NoPE more formally: they show it can in principle represent both absolute and relative position, and — the result that surprises most people — in their experiments NoPE outperformed the explicit schemes they compared (learned absolute, sinusoidal, ALiBi, RoPE, T5-style relative bias) at generalizing to sequences longer than training. Their conclusion is that “explicit position embeddings are not essential for decoder-only Transformers to generalize well to longer sequences.”
Read the scope carefully before over-updating on that: those are small decoder-only models trained on downstream reasoning tasks, not frontier-scale language models on open text, and the field did not switch. Nobody has published a replication at 70B scale, and the absence of NoPE frontier models is weak evidence rather than a refutation.
What NoPE is genuinely useful for is as a sanity check on what position encodings are actually doing. They’re not creating positional information out of thin air — the causal mask already leaks some. They’re giving the model a more direct, less learning-required handle on position than it would otherwise have to discover.
Why long-context is mostly a RoPE-extension story
If you train a RoPE model on sequences up to 4k tokens, the rotation angles m · θ_k for the slowest-rotating pairs sweep through some specific range of angles for m ∈ [0, 4k]. At inference time, if you suddenly feed it a 32k-token sequence, those slow-rotating pairs see angles 8× larger than anything in training. The attention math doesn’t crash, but the model has never seen that part of the rotation phase space. Behavior degrades.
The fixes are all variants of “rescale θ so the angles stay in the trained range”:
- Position interpolation (PI) — Chen et al. (2023): squash the new positions back into the training range by scaling positions down (i.e. divide m by the extension factor). Cheap, requires fine-tuning.
- NTK-aware scaling — community-developed (the bloke / kaiokendev): instead of scaling positions uniformly, rescale the base (the 10000) so high-frequency dimensions are barely changed and low-frequency ones are scaled more. Often works without fine-tuning.
- YaRN — Peng, Quesnelle, Fan, Shippole (arXiv 2309.00071, 2023): a more careful per-dimension scaling plus an attention-temperature adjustment. The headline claim from the paper is 10× fewer tokens and 2.5× fewer training steps than prior context-extension methods to reach the same quality. (YaRN turns up often in open-weights long-context fine-tunes, but nobody has surveyed how common it actually is, so treat prevalence claims about it as impressions.)
All of these are RoPE-specific. They’re not adjustments to attention, or to the KV cache, or to the model weights — they’re knobs on the rotation schedule. That’s the deepest sense in which the choice of positional encoding shapes what the long-context regime even looks like.
The seams worth seeing
- Position is not the same as causality. The causal mask makes a model autoregressive; the positional encoding tells it where in the sequence each token is. Encoder-only models (BERT) need positional encoding without a causal mask. Decoder-only models need both, but as NoPE shows, the mask alone leaks some position info “for free.”
- RoPE is per-head and per-pair. Different attention heads in the same model are looking at different rotation frequencies for different pairs of dimensions. There’s no one place where “the position” lives in the activations; it’s distributed across rotation phases.
- There’s nothing magical about 10000. It’s a hyperparameter, picked once in 2017 for sinusoidal PE and inherited into RoPE. Increasing the base is one of the standard knobs (it’s roughly what NTK scaling is doing). Some recent models train with much larger bases on purpose to get better long-context out of the box.
- “Position” is positions of tokens, not positions of characters or bytes. If your tokenizer splits a word weirdly, that’s the granularity of the position signal. Position encoding sits on top of tokenization; it doesn’t fix tokenization’s failure modes.
You started with positional encoding = a position-dependent signal injected so attention can tell tokens apart. What did this post add? — + a choice about whether that signal is absolute or relative, and that choice is the whole long-context story. Telling “the dog bit the man” from “the man bit the dog” needs only some order signal; keeping that ability at token 500,000 needs a signal whose math depends on the gap between tokens rather than on their index, which is why the field’s long-context knobs are all knobs on a rotation schedule.
Check yourself
Before you go — you take a model trained with 4k-token RoPE and feed it a 40k-token document, with no rescaling. It doesn’t crash and it doesn’t output garbage; it just gets noticeably worse at questions about the middle of the document. Using the rotation picture, why would that be the failure shape rather than a hard error?
Answer
Nothing in the math breaks at position 40,000 — m · θ_k is just a bigger angle, and rotation is defined for any angle. What breaks is familiarity: the slow-rotating dimension pairs are now sweeping through phase ranges the model never saw a single training example in, so the attention scores they produce aren’t calibrated against anything learned. That degrades quality smoothly instead of erroring. It’s also why the fixes (position interpolation, NTK-aware scaling, YaRN) are all reparameterizations of θ that squeeze inference-time angles back into the trained range, rather than changes to attention itself.
And one more: an encoder-only model like BERT cannot drop its positional encoding, but NoPE decoder-only models train fine without one. What single architectural difference explains that?
Answer
The causal mask. In a decoder-only model, token i can only attend to tokens 1…i, so the number of things visible is itself a position signal — index 5 sees five predecessors, index 50 sees fifty — and Haviv et al.’s probes find models pick up an implicit notion of position from it. BERT has no mask; every token sees every other one, so the input really is an unordered set and “the dog bit the man” is indistinguishable from “the man bit the dog” without an explicit signal. This also reframes what a positional encoding buys in the decoder case: not information from nothing, but a direct handle instead of one the model has to reconstruct.
Famous related terms
- Sinusoidal positional encoding —
sinusoidal PE = sin/cos at geometric frequencies, added to input embeddings. The 2017 original. Largely retired in favor of RoPE for new models. - Learned absolute PE —
learned PE = lookup table of position → vector, added to input embeddings. BERT-era choice. Zero extrapolation past training length; you literally have no vector for unseen positions. - RoPE (Rotary Position Embedding) —
RoPE = rotate Q and K by an angle proportional to position, applied per layer. Makes the positional part of attention scores depend only on relative offset. The default across the open-weights families whose configs you can read. - ALiBi —
ALiBi = subtract a per-head linear penalty proportional to |i − j| from attention scores. No PE on Q/K; bias the scores instead. Better extrapolation than sinusoidal in the original paper; mostly bypassed by the RoPE-extension ecosystem. - NoPE —
NoPE = no explicit position signal; let the causal mask leak it. Only possible in decoder-only models. Beat explicit schemes at length generalization in Kazemnejad et al.’s experiments — on small models and reasoning tasks, and nobody has shown it holds at frontier scale. - YaRN —
YaRN = RoPE + per-dimension scaling + attention temperature adjustment. A recipe for taking a 4k-trained model to 32k+ without redoing pretraining; its paper reports reaching the same quality with 10× fewer tokens than earlier extension methods. - Position interpolation —
PI = scale positions down so they fit into the trained rotation range. Simpler than YaRN, usually needs fine-tuning.
Going deeper
- Su et al., RoFormer: Enhanced Transformer with Rotary Position Embedding (2021) — go here for the actual derivation of why rotating Q and K makes the positional term depend only on
m − n. - EleutherAI blog, Rotary Embeddings: A Relative Revolution — read this if the rotation argument didn’t click; it’s the clearest plain-English walk from “why not just add a vector” to RoPE.
- Kazemnejad et al., The Impact of Positional Encoding on Length Generalization in Transformers (NeurIPS 2023) — the rabbit hole: a head-to-head on which encoding schemes actually survive being asked about positions they never trained on.