Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why do positional encodings exist?

A transformer cannot tell 'dog bites man' from 'man bites dog' on its own. The attention math is symmetric in token order — until you bolt on a position signal. Every modern LLM does, and the choice of how shapes long-context behavior more than people realize.

AI & ML intermediate Apr 29, 2026 · updated Aug 25, 2026 · 18 min read

On this page

The picture version

Seven pictures for a reader who has never thought about word order as a problem. The prose below fills in the seams the pictures skip.

1 · The problem

Same six words. Attention cannot tell them apart.

the dog bit the man the man bit the dog two very different sentences the only thing separating them is order what attention receives { the, the, dog, bit, man } a set. sets have no order. both sentences → the same set Attention is a weighted sum over (key, value) pairs. nothing in that formula says where any token sits so somebody had to add the ability to tell those two sentences apart
The property has a name: permutation equivariant — shuffle the input and the output shuffles to match, with no order privileged. Exactly what you don’t want for language, where order is most of the meaning.

2 · The first fix

Give every token a numbered jersey before it walks in.

token vector + position #4 sin / cos, fixed into layer 1 one vector per position, the same one every time low dimensions oscillate fast, high dimensions oscillate slowly It works — and it encodes the wrong thing. this is absolute position: “you are token 47”. what attention usually wants is “the token three to my left”.
The paper hoped sinusoids would let models extrapolate past their training length. Later measurement didn’t bear that out — when Press et al. tested it directly, sinusoidal models degraded past their training length rather than extending gracefully. BERT-style learned position tables have no extrapolation property at all: positions past training are literally untrained vectors.

3 · The move that won

Don’t label the token. Rotate it.

position m position n m − n each Q and K vector is cut into 2D pairs and each pair is turned by an angle proportional to its position R(α) · R(β)ᵀ = R(α − β) so the positional part of the score depends only on the gap, never on the two indices “bit” attending to “dog” one token to its left scores the same whether the phrase starts at token 1 or token 40,000 and still differs from one token to its right which is the whole difference between our two sentences
No new parameters — it is a fixed transform on Q and K. Position is reinjected at every layer rather than added once at the bottom, and the rotated key is what gets cached, so it composes with the KV cache. Su et al. also show a mild built-in decay of attention with distance.

4 · Two roads not taken

Bias the scores instead — or encode nothing at all.

ALiBi — penalise distance don’t touch Q or K at all near far subtract a per-head slope times the gap NoPE — encode nothing the causal mask already leaks position token 1 sees 1 thing token 4 sees 4 things decoder-only models only — BERT can’t do this NoPE beat every explicit scheme at length generalisation — in one study. on small decoder-only models and reasoning tasks, not frontier-scale language models; nobody has shown it holds at 70B, and the field did not switch
NoPE’s real value is as a sanity check on what positional encodings do: they aren’t creating position out of thin air, since the mask already leaks some. They hand the model a direct handle instead of one it must reconstruct.

5 · Why long context breaks

The angles wander into phase space the model never trained on.

trained: positions 0 → 4k a small arc of the circle asked for: positions 0 → 32k eight times further round Nothing crashes. Rotation is defined for any angle. so quality degrades smoothly instead of erroring — which is exactly why it is so easy to ship without noticing
What breaks is familiarity, not arithmetic: the slow-rotating pairs now sweep through phases no training example covered, so the scores they produce aren’t calibrated against anything learned. Every context-extension recipe is a way of squeezing those angles back into the trained range.

6 · The knobs

Every long-context recipe is a knob on the rotation schedule.

position interpolation divide every position by the extension factor squash them back in cheap; needs fine-tuning NTK-aware scaling rescale the frequency base, not the positions fast pairs barely touched often works without fine-tuning YaRN per-dimension scaling plus an attention-temperature tweak the careful version its paper reports 10× fewer tokens None of these touch attention, the weights, or the cache. they are all reparameterisations of one angle schedule — which is the deepest sense in which this choice shapes the long-context regime
All three are RoPE-specific, and that is the point. Pick a positional scheme and you have picked what your long-context story is going to look like — the entire extension literature exists because the field settled on rotations.

7 · Keep this card

The whole thing on one index card.

positional encoding = a position-dependent signal injected before the attention scores are computed + a choice: absolute or relative ∴ that second line is the whole long-context story
Picture to keep: every token walks into the room wearing a numbered jersey — and in the modern version the number isn’t printed on the jersey at all, it’s the angle each player is turned to face, so what attention reads off is how far apart two players are pointing. Where it breaks: the rotation is applied to dozens of 2D slices at different speeds, so there is no single “position” written down anywhere.

Why it exists

Type “the dog bit the man” into any chatbot, then type “the man bit the dog”, and you get two different answers — obviously, because they mean different things. Nothing about that feels like an achievement. But those two prompts are the same six words. The only thing separating them is order, and the core operation inside a transformer is mathematically blind to order. Somebody had to add the ability to tell those two sentences apart.

That’s the embarrassing fact the original transformer paper had to fix in a footnote-shaped way: attention has no idea what order its input tokens came in.

Take that sentence, shuffle the tokens, and feed both versions through pure self-attention: the math comes out the same up to a permutation of the outputs. Attention is a weighted sum over a set of (key, value) pairs. Sets have no order. “the dog bit the man” and “the man bit the dog” produce the same set, so the same attention scores, so the same internal representations — just rearranged. That property has a name: permutation equivariant. It’s exactly what you don’t want for language, where order is most of the meaning.

The original Attention Is All You Need paper (Vaswani et al., 2017) handled this by adding a fixed sinusoidal vector to each token’s input embedding — one vector per position, the same one every time, baked from sines and cosines at geometrically-spaced frequencies. Token embedding plus position vector goes into the model. That’s the entire trick.

That early choice has aged into a much bigger story. The open-weights LLM families whose architectures are public — LLaMA, Mistral, Gemma, Qwen, GPT-NeoX — don’t use the original sinusoidal scheme. They use RoPE, which works by rotating the query and key vectors inside the attention layer rather than adding anything to the embeddings. (Closed frontier models mostly don’t document their positional scheme, so treat “everyone uses RoPE” as an inference from the open ecosystem, not a surveyed fact.) The reason matters, and it’s why “long context” was hard in a way nobody quite expected.

Why it matters now

Three things make positional encoding a load-bearing choice today, not a footnote:

The choice of positional encoding is one of those design decisions where the wrong answer doesn’t fail loudly. It just makes the model quietly worse at long-range work.

The short answer

positional encoding = a position-dependent signal injected into Q/K (or input embeddings) so attention can tell tokens apart by where they are

Picture to keep: every token walks into the room wearing a numbered jersey — and in the modern version the number isn’t printed on the jersey at all, it’s the angle each player is turned to face, so what attention reads off is how far apart two players are pointing. Where the analogy breaks: a jersey number is one label per player, while the rotation is applied separately to dozens of 2D slices of each vector at different speeds, so there’s no single “position” written down anywhere.

Self-attention by itself sees a bag of tokens, not a sequence. To get sequence behavior you have to inject “you are token #i” somewhere in the pipeline before the attention scores get computed — into the input embeddings, or into the Q and K vectors that attention actually multiplies. The original transformer added a fixed sinusoid to each input embedding. Modern open-weights LLMs (LLaMA, Mistral, GPT-NeoX, Gemma, Qwen) instead rotate the query and key vectors inside attention — that’s RoPE — because it makes the positional part of the score depend only on the relative distance between two tokens, which generalizes and composes much better.

How it works

Start with the failure mode, because the rest follows from it.

What “permutation equivariant” actually buys you

A single attention head computes, for each token i:

output_i = Σ_j  softmax_j( q_i · k_j / √d ) · v_j

Notice what this depends on: the content of q_i, and the contents of all k_j and v_j. There is nothing here about where token j sits in the sequence. If you shuffle the tokens, you shuffle the (k, v) pairs, but the set of pairs is unchanged — so for any given query, the scores it computes are unchanged (modulo which output slot you read out).

For a vision model on patches that’s sometimes fine; spatial position can be encoded in other ways. For language, it’s catastrophic: “the dog bit the man”, “the man bit the dog”, and “bit the the man dog” are all the same multiset of tokens.

You have to break the symmetry. The question is where you break it, and how.

Option 1: Add position to the input (sinusoidal, learned absolute)

The original transformer’s move: precompute a vector PE_i for each position i, and add it to the token embedding before layer 1.

x_i = embedding(token_i) + PE_i

Vaswani et al. defined PE_i using sines and cosines:

PE_{i, 2k}     = sin(i / 10000^{2k/d})
PE_{i, 2k+1}   = cos(i / 10000^{2k/d})

The 10000 is just a chosen base; the shape is geometric: low dimensions oscillate fast, high dimensions oscillate slowly. The reason for sines and cosines specifically (rather than learning the position vectors) was, in their own words, that it “may allow the model to extrapolate to sequence lengths longer than the ones encountered during training.” A nice hope, and one the later literature didn’t bear out: when Press et al. measured extrapolation directly for the ALiBi paper, sinusoidal models degraded past their training length rather than gracefully extending. The construction stuck for years anyway.

BERT and many follow-ups used learned absolute position embeddings instead: a lookup table from position index to a learned vector, just like the token embedding table. Same shape, simpler, no extrapolation property at all (positions past training length are literally untrained vectors).

The intuitive worry with both flavors of “add to input”: the position information enters once, at the bottom, and has to survive through every layer’s mixing. By layer 24 the model has been shuffling these representations around for a long time. Whether the signal is meaningfully attenuated in practice isn’t settled in the literature — treat it as motivation for what comes next, not as established fact. What is established is that absolute schemes encode absolute position, when what attention usually wants is relative position — token i caring about “the token three to my left” rather than “the token at index 47.”

Option 2: Rotate Q and K (RoPE)

RoPE — proposed in Su et al.’s RoFormer paper (Su, Lu, Pan, Murtadha, Wen, Liu, 2021; arXiv 2104.09864) — does something cleaner. Instead of adding to the embedding, it rotates the query and key vectors themselves, by an amount that depends on the position.

The key fact, written compactly: split each head’s d-dimensional Q and K into d/2 pairs of components. For each pair, treat it as a 2D vector and rotate it by an angle m · θ_k, where m is the position and θ_k is a per-pair frequency (using the same 10000-base geometric progression as the original sinusoidal scheme). Different pairs rotate at different rates. The whole operation is a 2×2 rotation applied independently to each pair.

The magic: the dot product of a rotated query at position m with a rotated key at position n depends on the content of q and k plus the relative offset m − n, not on the absolute m and n separately. This is a mathematical fact about rotations: R(α) · R(β)ᵀ = R(α − β). Concretely: “bit” attending to “dog” one token to its left produces the same positional contribution whether that phrase starts at token 1 or token 40,000 — and it’s still a different contribution from “bit” attending to “dog” one token to its right, which is the whole difference between our two sentences. Translate the sequence and the attention pattern stays identical; flip the order and it doesn’t.

Other useful properties of RoPE:

LLaMA, Mistral, Gemma, GPT-NeoX, GPT-J, and Qwen all use RoPE — it’s the default across the open-weights families whose configs you can read. What the closed frontier labs use is mostly undocumented, so this is a claim about the open ecosystem rather than a surveyed share.

Option 3: Bias the attention scores directly (ALiBi)

A third route: don’t touch Q or K at all. Just bias the attention score for token i attending to token j by −|i − j| · m_h, where m_h is a small per-head slope. Closer tokens get a smaller penalty; farther tokens get a larger one. This is ALiBi, from Press, Smith, and Lewis (arXiv 2108.12409, ICLR 2022).

ALiBi’s pitch in the paper was specifically about extrapolation: train at length 1024, test at 2048+, with no fine-tuning. The headline result was a 1.3B model trained on 1024 tokens that extrapolated to 2048 with the same perplexity as a sinusoidal model trained at 2048, while training 11% faster and using 11% less memory. ALiBi has been used in some real models (e.g. MPT, BLOOM), but the open-weights families that came after mostly chose RoPE. One plausible read: RoPE’s relative-position-aware Q/K rotations look more like a general mechanism, where ALiBi is a hand-designed bias term that happens to work for distance-decay specifically. No public source attributes the convergence to a single reason, so treat that as a take rather than a finding.

Option 4: No positional encoding at all (NoPE)

This one is the most surprising. Haviv et al. (arXiv 2203.16634, 2022) showed that causal (decoder-only) transformers with no positional encoding at all are still competitive with explicit-PE ones at language modeling, across multiple datasets and model sizes. The intuition the paper develops is that the causal mask itself — token i can only attend to tokens 1…i — leaks position information: a token at position 5 has 5 things to attend to; a token at position 50 has 50. (Their probing experiments suggest the model picks up an implicit notion of absolute position; the exact mechanism is more subtle than a literal “count.”)

Kazemnejad et al. (The Impact of Positional Encoding on Length Generalization in Transformers, arXiv 2305.19466, 2023) studied NoPE more formally: they show it can in principle represent both absolute and relative position, and — the result that surprises most people — in their experiments NoPE outperformed the explicit schemes they compared (learned absolute, sinusoidal, ALiBi, RoPE, T5-style relative bias) at generalizing to sequences longer than training. Their conclusion is that “explicit position embeddings are not essential for decoder-only Transformers to generalize well to longer sequences.”

Read the scope carefully before over-updating on that: those are small decoder-only models trained on downstream reasoning tasks, not frontier-scale language models on open text, and the field did not switch. Nobody has published a replication at 70B scale, and the absence of NoPE frontier models is weak evidence rather than a refutation.

What NoPE is genuinely useful for is as a sanity check on what position encodings are actually doing. They’re not creating positional information out of thin air — the causal mask already leaks some. They’re giving the model a more direct, less learning-required handle on position than it would otherwise have to discover.

Why long-context is mostly a RoPE-extension story

If you train a RoPE model on sequences up to 4k tokens, the rotation angles m · θ_k for the slowest-rotating pairs sweep through some specific range of angles for m ∈ [0, 4k]. At inference time, if you suddenly feed it a 32k-token sequence, those slow-rotating pairs see angles 8× larger than anything in training. The attention math doesn’t crash, but the model has never seen that part of the rotation phase space. Behavior degrades.

The fixes are all variants of “rescale θ so the angles stay in the trained range”:

All of these are RoPE-specific. They’re not adjustments to attention, or to the KV cache, or to the model weights — they’re knobs on the rotation schedule. That’s the deepest sense in which the choice of positional encoding shapes what the long-context regime even looks like.

The seams worth seeing

You started with positional encoding = a position-dependent signal injected so attention can tell tokens apart. What did this post add? — + a choice about whether that signal is absolute or relative, and that choice is the whole long-context story. Telling “the dog bit the man” from “the man bit the dog” needs only some order signal; keeping that ability at token 500,000 needs a signal whose math depends on the gap between tokens rather than on their index, which is why the field’s long-context knobs are all knobs on a rotation schedule.

Check yourself

Before you go — you take a model trained with 4k-token RoPE and feed it a 40k-token document, with no rescaling. It doesn’t crash and it doesn’t output garbage; it just gets noticeably worse at questions about the middle of the document. Using the rotation picture, why would that be the failure shape rather than a hard error?

Answer

Nothing in the math breaks at position 40,000 — m · θ_k is just a bigger angle, and rotation is defined for any angle. What breaks is familiarity: the slow-rotating dimension pairs are now sweeping through phase ranges the model never saw a single training example in, so the attention scores they produce aren’t calibrated against anything learned. That degrades quality smoothly instead of erroring. It’s also why the fixes (position interpolation, NTK-aware scaling, YaRN) are all reparameterizations of θ that squeeze inference-time angles back into the trained range, rather than changes to attention itself.

And one more: an encoder-only model like BERT cannot drop its positional encoding, but NoPE decoder-only models train fine without one. What single architectural difference explains that?

Answer

The causal mask. In a decoder-only model, token i can only attend to tokens 1…i, so the number of things visible is itself a position signal — index 5 sees five predecessors, index 50 sees fifty — and Haviv et al.’s probes find models pick up an implicit notion of position from it. BERT has no mask; every token sees every other one, so the input really is an unordered set and “the dog bit the man” is indistinguishable from “the man bit the dog” without an explicit signal. This also reframes what a positional encoding buys in the decoder case: not information from nothing, but a direct handle instead of one the model has to reconstruct.

Going deeper