Why does cosine similarity dominate over Euclidean distance in embeddings?
Two vectors can be far apart and still mean the same thing. Cosine similarity asks the only question that turns out to matter: are they pointing the same way?
On this page
The picture version
Five pictures for a reader who has never compared two lists of numbers. The prose below fills in the seams the pictures skip.
1 · The problem
Two playlists, one ten times bigger, exactly the same taste.
2 · The obvious measure gets it wrong
Straight-line distance says they’re miles apart. It’s measuring size.
3 · The move
Draw them as arrows from the same point and measure the angle instead.
4 · Why that’s the right question for meaning
In a space of meanings, length mostly records how much text there was.
5 · Keep this card
The whole thing on one index card.
Why it exists
Picture two Spotify playlists. Yours has 200 songs; a friend’s has 20. Both are roughly 60% rock, 30% pop, 10% jazz. Do you have similar taste? If you measured “how far apart are the raw song counts,” you’d say no — yours is ten times bigger. If you measured “are the proportions the same,” you’d say yes. Cosine similarity is the second measurement. It throws away how much there is and only asks whether two things point in the same direction. That’s exactly the question you want to ask of an embedding, and exactly the wrong question to ask of a map.
If you’ve ever wired up RAG, opened a vector database, or read the docs for any embedding API, you’ve seen the same line: use cosine similarity to compare vectors. Almost nobody stops to ask why. Euclidean distance — the straight-line ruler distance you learned in school — is right there, it works in any number of dimensions, and it’s what “distance” means in normal life. So why did the entire field quietly agree to ignore it?
The short version is: when you turn meaning into a vector, the direction of the vector is the part that carries meaning, and the length of the vector tends to drift around for boring reasons — how long the document was, how confident the model felt, how many tokens it averaged over. Euclidean distance treats length and direction as equally important. Cosine similarity throws length away on purpose. In embedding space, that turns out to be exactly the right thing to do.
This post is about why that works out, what cosine similarity actually is once you cut through the formula, and where it stops being the right tool.
Why it matters now
The math here is old and hasn’t changed — cosine of an angle is cosine of an
angle. What changed is how much software depends on it. Semantic search, RAG
retrieval, near-duplicate detection, recommendation candidate generation:
all of them are nearest-neighbor lookups in a vector space, run constantly by
people who never picked the metric. The choice of similarity metric is
buried inside pgvector, Pinecone, FAISS, Qdrant — and it shows up as a
configuration flag (cosine / dot / l2) that most teams set once and
never revisit.
Picking the wrong one quietly degrades retrieval quality. A RAG system that uses Euclidean distance over un-normalized embeddings will rank short, terse passages systematically differently from long, verbose ones — not because they mean different things but because their vectors are different lengths. The model never told you that was going to happen. The metric did.
So this is one of those small choices that sits under a lot of working software. Worth understanding once.
The short answer
cosine_similarity(a, b) = (a · b) / (‖a‖ · ‖b‖) = cos(θ)
Picture to keep: two arrows pinned at the same origin. Cosine similarity measures the angle between them and refuses to look at their lengths — your 200-song playlist and your friend’s 20-song playlist are the same arrow, one just drawn longer.
It’s the cosine of the angle between two vectors. It ignores how long each vector is and only asks how aligned they are. Two vectors pointing the same way score 1, perpendicular ones score 0, opposite ones score −1. Euclidean distance, by contrast, asks “how far apart are the two tips?” — which mixes “different direction” and “different length” into a single number. For text embeddings, direction is the part the training objective shaped; length mostly isn’t.
How it works
The naive attempt. Turn each playlist into a vector of genre counts —
yours is [120, 60, 20] (rock, pop, jazz), your friend’s is [12, 6, 2] —
and measure similarity the way you’d measure distance on a map:
Euclidean distance, ‖a − b‖, the straight line between the two tips.
Smaller means more similar. This is the definition of “distance” everyone
already owns, and it’s one of the two or three options every vector database
puts in front of you.
Why it breaks. ‖a − b‖ ≈ 122 here, which is enormous — and yet the two
playlists have literally identical taste. The distance is large entirely
because one list is ten times longer than the other. Euclidean distance folds
“different direction” and “different size” into one number, and in embedding
space size is the part you don’t care about.
The fix. Divide the size out before you compare: cosine similarity,
(a · b) / (‖a‖ ‖b‖). Normalize both vectors to unit length first, then take
the dot product. Larger is more similar. Geometrically, you slide both vectors
to the origin, project them onto the unit sphere, and then measure how close
they are. Length information is destroyed by the projection — deliberately.
Direction is all that’s left, and both playlists land on exactly the same
point: cos = 1.
Why direction is the meaningful part
Embedding models are trained with objectives like “pull these two representations together, push those two apart.” The training signal shapes where the vector points. It does not, in general, pin down a canonical length. Two side effects fall out of that:
- Vectors of related things end up clustered along similar directions — that’s the property the loss explicitly rewarded.
- Vector magnitudes pick up things you didn’t ask for — passage length
and token count being the usual suspects. Models often pool over tokens
(mean, max, last hidden state of a
[CLS]token), and that pooling step is a plausible route by which length leaks into magnitude. The mechanism is well-motivated rather than measured: no published figure quantifies how much of magnitude variance is length for a given model.
If you compare with Euclidean distance, point 2 contaminates point 1. A short query and a long document about the exact same topic can have very different magnitudes, so their tips can be far apart in raw space even though their directions agree. Cosine sidesteps the issue: project both onto the unit sphere, look at the angle, done.
The “they’re almost the same metric” trick
Here’s the part that’s worth carrying around in your head. If both
vectors have been normalized to unit length (‖a‖ = ‖b‖ = 1), then:
‖a − b‖² = ‖a‖² + ‖b‖² − 2(a · b)
= 1 + 1 − 2(a · b)
= 2 − 2 · cos(θ)
So on the unit sphere, Euclidean distance and cosine similarity are monotonically related — ranking by one gives the same nearest neighbors as ranking by the other. That’s why a vector database that “only supports L2” can still do cosine search: normalize the vectors at insert time and normalize the query the same way, and the index returns the cosine ranking. Forgetting to normalize the query is the classic version of this bug — the equivalence needs both sides on the unit sphere.
It’s also why, in practice, lots of systems quietly store unit-normalized embeddings and use plain dot product as the similarity score. With normalization, dot product is cosine. Without it, dot product is “cosine, weighted by how big the vectors happen to be” — sometimes useful (popular items get a magnitude boost in some recsys setups) but usually a footgun.
The numbers, side by side
Add a third listener to the playlist example — someone with a 15-song all-jazz playlist:
A = [120, 60, 20] # you: 200 songs, 60/30/10
B = [ 12, 6, 2] # your friend: 20 songs, identical proportions
C = [ 0, 0, 15] # a pure-jazz listener, 15 songs
Euclidean distances: ‖A − B‖ ≈ 122, ‖A − C‖ ≈ 134. Euclidean does rank B
closer than C — but only by about 9%. By that ruler, “identical taste” and
“nothing in common” are nearly the same answer.
Cosine similarities: cos(A, B) = 1, cos(A, C) ≈ 0.15. Identical direction
versus almost none. That’s the separation a search system actually needs.
Now imagine these vectors are 1536-dimensional embeddings, the magnitude gap is correlated with document length across your corpus, and you’re asking “rank ten thousand documents by how much they match this query.” When the meaningful signal is a few percent of the distance and length is most of the rest, the ranking degrades in a way nothing in your logs will announce.
Where cosine is the wrong tool
Cosine isn’t universally correct. A few honest exceptions:
- When magnitude is meaningful. If you’ve designed your vectors so that “more important” or “more confident” really does mean “longer,” cosine throws that away. Some classical TF-IDF and recsys setups intentionally lean on magnitude.
- When the embedding model wasn’t trained for it. Some image and recsys embeddings are trained against Euclidean / L2 objectives; using cosine on those is a metric mismatch, and you can usually tell by reading the model card. (My read is that the AI-era text embedding models have converged hard on cosine/dot, while the broader ML world has not — though no survey puts a number on the split.)
- For dense numerical features (sensor readings, geographic coordinates, anything where the axes have units). Cosine is mostly a semantic-space tool. On a map, Euclidean is what you want.
You started with cosine_similarity = (a · b) / (‖a‖ ‖b‖). What did this post
add? — + the assumption that length is noise. That assumption is what makes
cosine right for embeddings and wrong for maps, and it’s the one line of the
formula that isn’t math: it’s a claim about your data that you have to check.
Famous related terms
- Dot product —
a · b = Σ aᵢ bᵢ— cosine’s unnormalized cousin. Same ranking as cosine when vectors are unit-length, faster to compute, and the primitive many vector indexes reduce to once vectors are normalized. - Euclidean distance (L2) —
‖a − b‖— straight-line distance. The default in geometry, the wrong default in semantic search unless you normalize first. - L2 normalization —
a / ‖a‖— projecting a vector onto the unit sphere. The bridge between “I have a Euclidean index” and “I want cosine semantics.” - ANN index — the data structure (HNSW, IVF, etc.) that makes cosine/dot/L2 search sub-linear at billion-vector scale.
- Embedding — see embeddings — the reason you ever need a similarity metric in the first place.
Going deeper
- The
pgvectorREADME’s operator table — for “what does my database actually compute when I pickcosinevsl2vsip?”, straight from the implementation rather than a tutorial. - 3Blue1Brown, “Dot products and duality”
— for “what is a dot product actually measuring?”, which is the geometric
intuition the identity
‖a − b‖² = ‖a‖² + ‖b‖² − 2(a · b)formalizes. - Rabbit hole: the model card for whichever embedding you actually use — for “was this model trained for cosine at all?”, the one question this post can’t answer on your behalf.
A note on what I’m sure of and what I’m not. The mathematical claims here — the formulas, the unit-sphere equivalence, the direction-vs-magnitude framing — are standard. The empirical claim that modern text embedding models are trained such that direction carries the signal and magnitude carries noise is the consensus story across model cards and tutorials, but no single citation establishes it for every embedding model in use. If a specific embedding you’re using disagrees, the model card is the source of truth, not this post.