Why do embeddings exist?
Computers want numbers, but you also want 'cat' and 'kitten' to live next to each other. Embeddings are the trick that makes both true at once.
On this page
The picture version
The whole idea in six pictures, for a reader who has never heard the word “embedding.” The prose below fills in the seams the pictures skip.
1 · The problem
The words you type and the words that answer you have nothing in common.
2 · The old way
Give every word a number, and the numbers mean nothing.
3 · The key idea
Stop choosing the numbers. Let one rule move them.
4 · The thing the rule hides
Somebody still has to say which pairs "should" be close. Text says it for free.
5 · The shortcut
Do the expensive half once, before anyone asks.
6 · Keep this card
The whole thing on one index card.
Why it exists
You want to cancel a subscription, so you open the company’s help centre and type what’s actually in your head: “stop the bill.” The right article is called “How do I cancel my subscription?” — and it comes up first. Look at those two strings. They share no words at all. “Stop” isn’t “cancel,” “bill” isn’t “subscription,” and nothing about the letters connects them.
A search box that only matched words would have returned nothing here, and you’d be emailing support. Something in that box knows the two phrases are about the same thing, and it isn’t a synonym list somebody wrote for your exact wording.
The problem underneath: a neural network can’t read the word “bill.” It multiplies numbers. So text has to become numbers somewhere. The boring way — give every word an arbitrary integer ID — is fine for indexing and destroys everything interesting. In an ID-based world “cat” and “kitten” are exactly as far apart as “cat” and “thermodynamics,” because the IDs were assigned alphabetically or by frequency or by whim.
Embeddings exist to give you numbers that keep the meaning. Two demands at once:
- It’s a vector of numbers, so a model or a database can do math on it.
- Similar things land near each other, so distance stands in for “are these alike?”
The second demand is the one that pays. Once you have it, a whole family of problems — find the doc that answers this, suggest songs like this song, cluster these tickets, flag these two reviews as near-duplicates — collapses into one operation: find the nearest points. The idea usually credited underneath all of it is distributional semantics: if “cat” and “kitten” show up around the same other words, maybe they mean similar things. Hold on to one question while you read, because it’s the whole game: who decides what “similar” means? We’ll follow “stop the bill” until that has an answer.
Why it matters now
Two different things get called embeddings, and it’s worth separating them. Every LLM has an internal embedding table that turns token IDs into vectors before the first layer — you get that for free, and it isn’t what people mean by “using embeddings.” The other kind is a separate model whose whole output is one vector per input — a document, a chunk, a query, an image, and that’s the one you deliberately reach for:
- RAG — the standard way to answer questions from your own documents — is an embedding search step with an LLM bolted on the end. The model is the visible part; retrieval decides what it even gets to see.
- Semantic search in a product — the “stop the bill” box above — is embeddings plus a vector index.
- Recommendations, clustering, deduplication, and few-label classification are all routine once the data lives in a sensible vector space.
- Vector databases (Pinecone, Weaviate, Qdrant, Chroma, Postgres with
pgvector) are a product category built around one query — “give me the nearest vectors” — which only becomes a useful query once embeddings exist.
The engineer’s reason to care: an off-the-shelf embedding model plus a vector store buys you a lot of “AI features” without fine-tuning anything. That’s my opinion about cost-effectiveness rather than a measured claim — but the timeline isn’t opinion. word2vec is from 2013, well before anything anyone would call an LLM today.
The short answer
embedding = a learned function (text | image | thing) → ℝᵈ such that semantic similarity ≈ vector similarity
Picture to keep: a map where you don’t choose the coordinates — the content does. Help articles about cancelling drift into one neighbourhood, articles about password resets into another, and your typed question lands as a pin somewhere on that map. Search is “what’s near the pin.”
An embedding is a fixed-length list of floats — a few hundred to a couple of thousand is the range you’ll meet in practice, though check whatever model you’re using — produced by a model trained so that things meaning similar stuff come out pointing in similar directions. “Are these two alike?” becomes arithmetic.
How it works
Keep two questions apart — how you build the map, and how you use it. Both are easiest to see by failing your way forward.
Attempt 1: match the words. Search for “stop the bill” by looking for documents containing those words.
Why it breaks: zero overlap with “cancel your subscription.” Keyword search is excellent when the user already speaks your product’s vocabulary, and useless the moment they don’t — which is most of the time, since the people searching for how to cancel are precisely the people who never learned your terminology.
Attempt 2: write a synonym list. Map “stop” → “cancel,” “bill” → “subscription,” “invoice,” “charge.”
Why it breaks: somebody has to maintain it, forever, per language, per product, and it doesn’t compose. “Stop the bill” needs the phrase to mean cancellation; “stop” alone might be about pausing a video. You’re back to hand-writing the knowledge, which is the thing you were trying to avoid.
Attempt 3: number the words. Give every word an ID and let a model learn from those.
Why it breaks: the IDs are arbitrary. If “bill” is 4102 and “subscription” is 88,301, the gap between them is an artifact of the sorting, not a fact about English. There’s nothing to learn from; the model would have to memorize every pair separately, which is Attempt 2 with extra steps.
Fix: let the corpus assign the numbers. Instead of picking coordinates, learn them, using one rule with two halves:
Pull together representations of things that should be similar. Push apart representations of things that shouldn’t.
The trick is defining “should” without a human labelling every pair — and text
supplies it free. word2vec (Mikolov et al., 2013), in its skip-gram form, slides a window over a
huge corpus, treats each word’s neighbours as “should be close,” and treats
randomly-drawn other words as “should be far.” Because “cancel” and “stop” appear
around similar words, they end up in similar places without anyone saying so. The
famous party trick — vec("king") − vec("man") + vec("woman") ≈ vec("queen") —
fell out as a side effect. (How reliable that analogy really is has been argued
about ever since; the durable point is that usable linear structure appeared in
the geometry at all.)
Why it breaks: one vector per word. “Stop the bill” is three vectors, and averaging them throws away that this is a phrase about billing rather than a sentence containing the word “bill.” Worse, a static vector for “bill” has to serve the invoice, the legislation, and the duck.
Fix: embed the whole passage in context. Run the sentence through a transformer, which lets each word’s representation depend on its neighbours, and pool the result into a single vector. Train it on pairs that are known to mean the same thing — question/answer pairs, paraphrases, click logs — so “stop the bill” and “cancel your subscription” are pulled together explicitly. Sentence-BERT (2019) is the clearest early example of this shape; contrastive training on sentence pairs is still the common recipe behind the embedding models you’d call today, though the specific architectures and data have moved on a lot. Note what just happened to our question: “similar” now means whatever the training pairs said it means. Nobody defined it; somebody chose examples. The same recipe crosses modalities — CLIP trains an image encoder and a text encoder together on hundreds of millions of (image, caption) pairs, so pictures land near the words that describe them.
You will almost never train one of these. You pick one off the shelf and treat it as a black box.
Why it breaks: comparing your query against two million help articles, one at a time, per keystroke.
Fix: do the expensive half offline. Embed every document once, store the vectors, and at query time embed only the query and look up its neighbours — with an ANN index so “nearest 10 of 100 million” is milliseconds rather than a table scan:
# offline, once
for chunk in docs:
store(id=chunk.id, vector=embed(chunk.text), payload=chunk)
# online, per question
q_vec = embed("stop the bill")
hits = vector_db.search(q_vec, top_k=8)
prompt = SYSTEM + format(hits) + user_question
answer = llm(prompt)
That’s the skeleton of RAG. Production systems layer a lot on top — keyword search blended with vector search, a reranker over the top-k, query rewriting, chunking strategy — but the retrieve-then-prompt shape is this. The LLM looks like the clever part; the embedding step is why the right eight chunks were in the prompt at all.
Where the map picture stops being safe, and the things that bite people:
- Direction, not distance. Cosine similarity is usually the right metric because these models encode meaning in the vector’s direction and let its magnitude drift. Two vectors pointing the same way count as similar even if one is much longer — which is not how an ordinary map works.
- Dimensionality is a budget, not a quality knob. 1536 dimensions aren’t “smarter” than 384 in any deep sense; they’re more expensive to store and search, and sometimes — not always — a bit more accurate. Matryoshka-style models let you truncate a vector to a shorter prefix and keep most of the quality, turning it into a slider.
- Different models’ spaces don’t line up. A vector from one model is meaningless to another. Swap embedding models and you must re-embed the corpus and the queries. In my experience this is the most common operational footgun there is.
- It encodes whatever the objective rewarded — usually topical similarity, but it can quietly pick up length, language, or formality too. When retrieval surprises you, suspect the embedding first.
- It goes stale. A model trained in 2022 has never seen a product launched in 2025 and can place it confidently in the wrong neighbourhood.
You started with embedding = learned function → vectors where similar means near. What did “stop the bill” add that the definition leaves out? — you never
specify what “similar” means; the training pairs do. That’s the whole reason it
scales past the synonym list, and also the reason it fails in ways you didn’t
choose: change the pairs and you change the geometry, and nobody wrote down what
the geometry now believes.
Famous related terms
- Vector / vector space —
vector = ordered list of numbers— the container; an embedding is a vector whose coordinates were learned. - Cosine similarity —
cos(a, b) = (a·b) / (‖a‖‖b‖)— the cosine of the angle between two vectors, and the default “are these alike?” function. - Vector database —
vector DB = store + ANN index over embeddings— built so “nearest 10 of 100M” returns in milliseconds. - ANN —
ANN ≈ nearest-neighbour search + an index allowed to be slightly wrong— without the “slightly wrong,” vector search at scale is unaffordable. - RAG —
RAG = embedding-based retrieval + LLM generation— the most common place embeddings show up in production code. - word2vec —
word2vec = shallow net + a context-prediction objective— the result that made people believe the geometry was real. - CLIP —
CLIP = image encoder + text encoder + contrastive (image, caption) loss— the recipe behind shared image/text spaces. - Tokenization — the step before embedding: chopping text into the discrete units the embedding table is indexed by.
Going deeper
- Mikolov et al., Efficient Estimation of Word Representations in Vector Space (2013) — the primary source for how you get a meaningful vector out of nothing but raw text and a sliding window.
- Reimers & Gurevych, Sentence-BERT (2019) — answers “why can’t I just average word vectors for a sentence?”, which is the question everyone hits second.
- The
sentence-transformersdocumentation — the hands-on explainer: read the “usage” pages to find out what actually happens when you embed your own corpus, including the choices (pooling, normalization, chunk size) this post glossed over. - Radford et al., Learning Transferable Visual Models From Natural Language Supervision (2021) — the rabbit hole: what happens when the same contrastive trick is pointed at images and captions at once.
A note on what I’m sure of: the high-level story — objectives, use cases, the contrastive shape, the “similar things point the same way” property — is well established. Benchmark numbers and “best embedding model right now” rankings change every few months; verify those against a current leaderboard rather than memorizing anything here.