Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why does cosine similarity dominate over Euclidean distance in embeddings?

Two vectors can be far apart and still mean the same thing. Cosine similarity asks the only question that turns out to matter: are they pointing the same way?

Math intro Apr 29, 2026 · updated Aug 25, 2026 · 10 min read

On this page

The picture version

Five pictures for a reader who has never compared two lists of numbers. The prose below fills in the seams the pictures skip.

1 · The problem

Two playlists, one ten times bigger, exactly the same taste.

your playlist — 200 songs rockpopjazz a friend’s — 20 songs rockpopjazz Same 60/30/10 split. Wildly different sizes. Similar taste, or not?
Both playlists are roughly 60% rock, 30% pop, 10% jazz — one is simply ten times longer. Whether they count as similar depends entirely on which question you ask, and the two obvious questions give opposite answers.

2 · The obvious measure gets it wrong

Straight-line distance says they’re miles apart. It’s measuring size.

rock count → the 20-song list the 200-song list a long way apart And that gap is almost entirely “one of them is bigger.”
Treated as points and measured with a ruler, the two lists are far apart — but nearly all of that distance is the size difference. The measure answers a question nobody asked: how much is there, rather than what kind of thing it is.

3 · The move

Draw them as arrows from the same point and measure the angle instead.

the angle between them: almost nothing the long arrow is just the same direction, drawn further One number, between −1 and 1, that refuses to look at length.
Pin both at the same origin and the size difference becomes arrow length, which the angle ignores entirely. Two things with the same proportions point the same way no matter how big they are — which is the question the playlists needed answered.

4 · Why that’s the right question for meaning

In a space of meanings, length mostly records how much text there was.

a one-line note on cancelling a three-page article on cancelling same direction — same subject different length — different amount of text So ignoring length is not throwing away information. It is throwing away the wrong information.
When text is turned into numbers, the direction carries what it is about and the length largely reflects how much of it there was. Discarding length is what makes a short note and a long article about the same topic register as similar.

5 · Keep this card

The whole thing on one index card.

cosine similarity = the angle between two arrows — and a refusal to look at their lengths Which is exactly wrong whenever the size is the thing you care about. two shopping baskets with the same proportions but very different totals are not the same customer, and a measure that throws away the total will tell you they are
Picture to keep: two arrows pinned at the same origin. It measures the angle between them and refuses to look at their lengths — your 200-song playlist and your friend’s 20-song playlist are the same arrow, one just drawn longer. Where that becomes the wrong tool: whenever the magnitude carries the meaning you were actually after.

Why it exists

Picture two Spotify playlists. Yours has 200 songs; a friend’s has 20. Both are roughly 60% rock, 30% pop, 10% jazz. Do you have similar taste? If you measured “how far apart are the raw song counts,” you’d say no — yours is ten times bigger. If you measured “are the proportions the same,” you’d say yes. Cosine similarity is the second measurement. It throws away how much there is and only asks whether two things point in the same direction. That’s exactly the question you want to ask of an embedding, and exactly the wrong question to ask of a map.

If you’ve ever wired up RAG, opened a vector database, or read the docs for any embedding API, you’ve seen the same line: use cosine similarity to compare vectors. Almost nobody stops to ask why. Euclidean distance — the straight-line ruler distance you learned in school — is right there, it works in any number of dimensions, and it’s what “distance” means in normal life. So why did the entire field quietly agree to ignore it?

The short version is: when you turn meaning into a vector, the direction of the vector is the part that carries meaning, and the length of the vector tends to drift around for boring reasons — how long the document was, how confident the model felt, how many tokens it averaged over. Euclidean distance treats length and direction as equally important. Cosine similarity throws length away on purpose. In embedding space, that turns out to be exactly the right thing to do.

This post is about why that works out, what cosine similarity actually is once you cut through the formula, and where it stops being the right tool.

Why it matters now

The math here is old and hasn’t changed — cosine of an angle is cosine of an angle. What changed is how much software depends on it. Semantic search, RAG retrieval, near-duplicate detection, recommendation candidate generation: all of them are nearest-neighbor lookups in a vector space, run constantly by people who never picked the metric. The choice of similarity metric is buried inside pgvector, Pinecone, FAISS, Qdrant — and it shows up as a configuration flag (cosine / dot / l2) that most teams set once and never revisit.

Picking the wrong one quietly degrades retrieval quality. A RAG system that uses Euclidean distance over un-normalized embeddings will rank short, terse passages systematically differently from long, verbose ones — not because they mean different things but because their vectors are different lengths. The model never told you that was going to happen. The metric did.

So this is one of those small choices that sits under a lot of working software. Worth understanding once.

The short answer

cosine_similarity(a, b) = (a · b) / (‖a‖ · ‖b‖) = cos(θ)

Picture to keep: two arrows pinned at the same origin. Cosine similarity measures the angle between them and refuses to look at their lengths — your 200-song playlist and your friend’s 20-song playlist are the same arrow, one just drawn longer.

It’s the cosine of the angle between two vectors. It ignores how long each vector is and only asks how aligned they are. Two vectors pointing the same way score 1, perpendicular ones score 0, opposite ones score −1. Euclidean distance, by contrast, asks “how far apart are the two tips?” — which mixes “different direction” and “different length” into a single number. For text embeddings, direction is the part the training objective shaped; length mostly isn’t.

How it works

The naive attempt. Turn each playlist into a vector of genre counts — yours is [120, 60, 20] (rock, pop, jazz), your friend’s is [12, 6, 2] — and measure similarity the way you’d measure distance on a map: Euclidean distance, ‖a − b‖, the straight line between the two tips. Smaller means more similar. This is the definition of “distance” everyone already owns, and it’s one of the two or three options every vector database puts in front of you.

Why it breaks. ‖a − b‖ ≈ 122 here, which is enormous — and yet the two playlists have literally identical taste. The distance is large entirely because one list is ten times longer than the other. Euclidean distance folds “different direction” and “different size” into one number, and in embedding space size is the part you don’t care about.

The fix. Divide the size out before you compare: cosine similarity, (a · b) / (‖a‖ ‖b‖). Normalize both vectors to unit length first, then take the dot product. Larger is more similar. Geometrically, you slide both vectors to the origin, project them onto the unit sphere, and then measure how close they are. Length information is destroyed by the projection — deliberately. Direction is all that’s left, and both playlists land on exactly the same point: cos = 1.

Why direction is the meaningful part

Embedding models are trained with objectives like “pull these two representations together, push those two apart.” The training signal shapes where the vector points. It does not, in general, pin down a canonical length. Two side effects fall out of that:

  1. Vectors of related things end up clustered along similar directions — that’s the property the loss explicitly rewarded.
  2. Vector magnitudes pick up things you didn’t ask for — passage length and token count being the usual suspects. Models often pool over tokens (mean, max, last hidden state of a [CLS] token), and that pooling step is a plausible route by which length leaks into magnitude. The mechanism is well-motivated rather than measured: no published figure quantifies how much of magnitude variance is length for a given model.

If you compare with Euclidean distance, point 2 contaminates point 1. A short query and a long document about the exact same topic can have very different magnitudes, so their tips can be far apart in raw space even though their directions agree. Cosine sidesteps the issue: project both onto the unit sphere, look at the angle, done.

The “they’re almost the same metric” trick

Here’s the part that’s worth carrying around in your head. If both vectors have been normalized to unit length (‖a‖ = ‖b‖ = 1), then:

‖a − b‖² = ‖a‖² + ‖b‖² − 2(a · b)
         = 1 + 1 − 2(a · b)
         = 2 − 2 · cos(θ)

So on the unit sphere, Euclidean distance and cosine similarity are monotonically related — ranking by one gives the same nearest neighbors as ranking by the other. That’s why a vector database that “only supports L2” can still do cosine search: normalize the vectors at insert time and normalize the query the same way, and the index returns the cosine ranking. Forgetting to normalize the query is the classic version of this bug — the equivalence needs both sides on the unit sphere.

It’s also why, in practice, lots of systems quietly store unit-normalized embeddings and use plain dot product as the similarity score. With normalization, dot product is cosine. Without it, dot product is “cosine, weighted by how big the vectors happen to be” — sometimes useful (popular items get a magnitude boost in some recsys setups) but usually a footgun.

The numbers, side by side

Add a third listener to the playlist example — someone with a 15-song all-jazz playlist:

A = [120, 60, 20]   # you: 200 songs, 60/30/10
B = [ 12,  6,  2]   # your friend: 20 songs, identical proportions
C = [  0,  0, 15]   # a pure-jazz listener, 15 songs

Euclidean distances: ‖A − B‖ ≈ 122, ‖A − C‖ ≈ 134. Euclidean does rank B closer than C — but only by about 9%. By that ruler, “identical taste” and “nothing in common” are nearly the same answer.

Cosine similarities: cos(A, B) = 1, cos(A, C) ≈ 0.15. Identical direction versus almost none. That’s the separation a search system actually needs.

Now imagine these vectors are 1536-dimensional embeddings, the magnitude gap is correlated with document length across your corpus, and you’re asking “rank ten thousand documents by how much they match this query.” When the meaningful signal is a few percent of the distance and length is most of the rest, the ranking degrades in a way nothing in your logs will announce.

Where cosine is the wrong tool

Cosine isn’t universally correct. A few honest exceptions:

You started with cosine_similarity = (a · b) / (‖a‖ ‖b‖). What did this post add? — + the assumption that length is noise. That assumption is what makes cosine right for embeddings and wrong for maps, and it’s the one line of the formula that isn’t math: it’s a claim about your data that you have to check.

Going deeper

A note on what I’m sure of and what I’m not. The mathematical claims here — the formulas, the unit-sphere equivalence, the direction-vs-magnitude framing — are standard. The empirical claim that modern text embedding models are trained such that direction carries the signal and magnitude carries noise is the consensus story across model cards and tutorials, but no single citation establishes it for every embedding model in use. If a specific embedding you’re using disagrees, the model card is the source of truth, not this post.