Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why RoPE replaced sinusoidal positional encoding

The original transformer added a fixed sine/cosine vector to each token. Almost no frontier model does that anymore. RoPE rotates queries and keys instead — and that one structural change is what made long context tractable.

AI & ML intermediate Apr 30, 2026 · updated Aug 25, 2026 · 12 min read

On this page

The picture version

Six pictures for a reader who has never thought about how a model knows word order. The prose below fills in the seams the pictures skip.

1 · The problem

Same six words, opposite meanings. By default the model can’t tell them apart.

the cat sat on the mat the mat sat on the cat what the attention layer receives { the, cat, sat, on, the, mat } A bag of words. Identical for both sentences. so every transformer has to bolt on some notion of where each word sits and how you bolt it on turns out to matter enormously once the text gets long
Attention compares every word with every other word, but nothing in that comparison knows which came first. Order has to be added deliberately — and this post is about the difference between two ways of adding it.

2 · The obvious fix

Give every slot a number, and add it to the word.

before the words enter the model at all cat cat + + slot 1 slot 5 = = one vector a different vector the two “cat”s now differ The symmetry is broken. The two sentences stop being the same input. job done, in the minimal sense — and this is what the original transformer did
Give slot 1 one vector, slot 5 another, and add it to the word before anything else happens. The same word in two places is now two different vectors, which is all you strictly needed to tell the two sentences apart.

3 · Why that turns out to be the wrong thing to encode

You told it which slot. What it actually needs is how far apart.

the same relationship, in two places in a long book the cat one word back slots 3 and 4, on page 1 the cat one word back slots 90,412 and 90,413 To attention these look like completely unrelated pairs of numbers. So “one word back” has to be learned again at every slot — and at slots it never saw in training, never at all. which is exactly the wall you hit asking about page 12 of a 300-page book
Whether one word follows another shouldn’t depend on whether the sentence starts on page 1 or page 200. Adding an absolute slot number encodes the wrong thing, and the model has to learn the same relationship over and over.

4 · The trick

Don’t add a number. Turn the vector, like a clock hand.

the word at slot m, turned by m steps the word at slot n, turned by n steps compare them: only the gap survives × = Turning by m, then comparing against something turned by n, leaves an answer that depends only on m − n. no learning required — it falls out of what rotation is
Turning preserves length and composes by adding angles, so when a word at one slot is compared against a word at another, the absolute slots cancel. What reaches the score is the distance between them — the property the last scene said the model was having to learn the hard way.

5 · Not one clock — dozens

Fast hands see the word next door. Slow hands see the other end of the book.

the vector is split into pairs, and each pair gets its own ticking rate whips round quick slow barely moves tells nearby words apart tracks the long haul One set of hands can register both “one word back” and “two hundred back.”
Each pair of numbers inside the vector turns at its own rate, from very fast to almost still. That spread is what lets a single mechanism register fine word order and coarse document position at the same time.

6 · Keep this card

The whole thing on one index card.

the change = stop adding a position to the word + start turning the word by its position ∴ only the gap between two words reaches the score position stops being a property of a word and becomes one of a pair And that is why stretching a model to a longer book became a knob you can turn: all the position information lives in one angle, so you rescale the angle rather than teach the model about slots it has never seen
Picture to keep: two clock hands — each word’s vector turned by an angle set by its position, so that comparing two of them leaves the angle between the hands, not either hand’s reading. Where it breaks: it isn’t one clock but dozens per vector, each ticking at its own rate.

Why it exists

Paste a 300-page book into a chat window and ask about something on page 12. Models do this now; a few years ago the request didn’t even typecheck, because context windows were a couple of thousand tokens. The thing that had to change to get from there to here isn’t mainly a bigger buffer — it’s how the model is told where each token sits.

Start with why it needs telling at all. “The cat sat on the mat” and “The mat sat on the cat” contain exactly the same six words, and mean different things. A transformer’s attention layer, by default, can’t tell them apart — it sees a bag of words, not an order. So every transformer has to bolt some notion of “position” onto its tokens before they can mean anything. The way you bolt it on turns out to matter enormously once the document gets long. The original transformer added a fixed sine/cosine pattern to each token. RoPE replaced that with a rotation — one of several things long context needed (the rest is attention kernels, memory management, and training data at length), but the one that turned “make this model read four times more” into a knob you can turn.

The original transformer, in other words, had a positional-encoding problem and a positional-encoding answer, and the answer turned out to be wrong in a way that only became obvious later.

Vaswani et al.’s Attention Is All You Need (2017) added a fixed vector — sines and cosines at geometrically-spaced frequencies — to each token’s input embedding. Position 0 got one vector, position 1 got another, and so on. The vector was glued onto the token before it ever entered the attention stack. The intuition was elegant: different frequencies encode different positional scales, and any relative offset between two positions can in principle be recovered as a linear function of those sines and cosines.

In practice, the scheme worked well enough to ship the original transformer, then was steadily replaced. Learned absolute embeddings (BERT-style) took over for a while. Relative position biases (T5-style) and later ALiBi went a different direction. And then in 2021 Jianlin Su and coauthors published RoPE in the RoFormer paper. Read the published configs of the open-weights families that followed — GPT-NeoX, LLaMA, Mistral, Falcon, Gemma, Qwen — and you find RoPE in all of them. Most were born with it rather than switched: the lineage just chose RoPE over sinusoidal from the start. (What the closed frontier models use isn’t documented; I’m describing the ecosystem you can inspect.)

The reason isn’t that sinusoidal “didn’t work.” It’s that adding position to the embedding is a mismatch with what you actually want, which is for attention scores between two tokens to turn on their relative distance rather than their absolute indices. RoPE encodes position as a rotation of the query and key vectors inside attention, and the relative-offset dependence then falls out of the dot product as an algebraic identity rather than something the model has to learn. That’s the case for it: not that addition fails, but that rotation gets the property for free.

Why it matters now

The switch isn’t just academic — it shows up in three places engineers hit constantly.

The short answer

RoPE = rotate Q and K by an angle proportional to position, before the dot product

Picture to keep: two clock hands. Each token’s query and key vectors get turned by an angle set by its position, and when attention multiplies a query at position m against a key at position n, what survives is the angle between the hands — the gap, not either hand’s absolute reading. Where the picture breaks: it isn’t one clock but dozens per vector, each ticking at its own rate, which is what lets the model see both “one word back” and “two hundred words back.”

Where sinusoidal positional encoding adds a position vector to each token embedding once at the input, RoPE rotates the query and key vectors at every attention layer by an angle that scales with the token’s position index. Because rotations preserve magnitude and compose by addition of angles, the positional part of the q·k score between a query at position m and a key at position n enters only through the difference m − n — the score as a whole still depends on what the two tokens mean, as it should. Position becomes a property of the interaction between two tokens, not a property baked into one of them.

How it works

Take the naive fix first, because RoPE is best understood as the repair for it.

Naive attempt: add a position vector to each token embedding. Give position 0 one vector, position 1 another, and add it to the token’s embedding before layer 1. This does break the symmetry — “cat” at position 1 is now a different vector from “cat” at position 5, so “the cat sat on the mat” and “the mat sat on the cat” stop being identical inputs. Job done, in the minimal sense.

Why it breaks. Two problems, and the second is the expensive one. First, position is now mixed into the same vector that carries meaning: the sum of a content vector and a position vector is one vector, and the model has to disentangle them. Second, and worse, what you encoded is each token’s absolute index — but attention almost always wants the relative one. Whether “cat” attends to “the” one token back shouldn’t depend on whether the phrase starts on page 1 or page 200. With absolute encodings the model has to learn that invariance separately at every offset it ever sees, and offsets it never saw during training it never learned at all. That’s precisely the wall you hit at page 12 of a 300-page book.

The fix. Pick a query vector q at position m and a key vector k at position n. RoPE splits each vector into 2-dimensional pairs and treats each pair as a point in the plane. For each pair i, it picks a frequency θᵢ (a geometric series, the same flavor sinusoidal used) and rotates the query pair by angle m·θᵢ and the key pair by n·θᵢ.

Because rotations in 2D compose by adding angles, the dot product of the two rotated pairs is:

rotated_q · rotated_k = |q||k| cos((m − n)·θᵢ + original_angle_between_q_and_k)

The absolute positions m and n dropped out; only their difference reaches the score, alongside the content term that was always there. Stack many such 2D pairs at different θᵢ and you’ve encoded position as a rotation pattern that the attention dot product naturally translates into a relative-distance signal. No addition, no learned table, no separate bias — just a structural choice about where in the embedding space “position” lives.

A few consequences fall out of this:

The honest seam: no clean ablation isolates RoPE versus sinusoidal at modern scale, controlling for everything else. The case for RoPE in the literature is partly mathematical (the relative-position property), partly empirical (RoFormer’s own results, then a cascade of large-model adoptions), and partly path-dependent — once LLaMA shipped with it, the open-source ecosystem standardized. Nobody has published a controlled head-to-head at, say, 70B parameters and 32k context. That result would be genuinely interesting.

You started with RoPE = rotate Q and K by an angle proportional to position. What did this post add that the definition doesn’t? — + position becomes a property of the *pair*, not of the token. That relocation is the whole story: because only m − n survives the dot product, the model learns “one token back” once instead of once per index, and stretching it to a 300-page book becomes a matter of rescaling an angle schedule rather than teaching it about positions it has never seen.

Check yourself

Before you go — sinusoidal encodings were advertised in the 2017 paper as extrapolating to unseen lengths, and RoPE makes no such promise; naive RoPE degrades past its training length too. So why did long-context work converge on RoPE anyway?

Answer

Because “degrades” is not the interesting property — “has a knob” is. RoPE concentrates all positional information in one place: the rotation angle m·θᵢ. That gives you a single parameterization to attack, which is exactly what position interpolation, NTK-aware scaling, and YaRN do — they rescale θ so a longer sequence’s angles land back in the range the model actually trained on. With additive encodings the position information is already summed into the embeddings, so there’s no equally clean place to intervene — and no comparable body of extension recipes was ever developed for them. RoPE won by being adjustable, not by being immune.

And one more: RoPE rotates queries and keys but leaves values alone. Does that mean values carry no positional information at all?

Answer

Not quite — it means values carry no directly injected position signal. Rotation happens on Q and K because their dot product is where position needs to show up: that product decides how much each value gets weighted. The values themselves are the content being mixed. But after one layer, each token’s representation is a position-weighted blend of other tokens’ values, so positional structure leaks into the representations that later layers turn into new values. The clean statement is that RoPE injects position into the attention scores; whatever positional structure the values carry is downstream of that.

Going deeper