Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why do embeddings exist?

Computers want numbers, but you also want 'cat' and 'kitten' to live next to each other. Embeddings are the trick that makes both true at once.

AI & ML intro Apr 29, 2026 · updated Aug 24, 2026 · 11 min read

On this page

The picture version

The whole idea in six pictures, for a reader who has never heard the word “embedding.” The prose below fills in the seams the pictures skip.

1 · The problem

The words you type and the words that answer you have nothing in common.

what you type stop the bill the article that answers you How do I cancel my subscription? must match stop bill cancel subscription not one shared word, not one shared letter pattern
A search box that only matched words returns nothing here. Something has to know these two phrases are about the same thing — and it can’t be a synonym list written for your exact wording.

2 · The old way

Give every word a number, and the numbers mean nothing.

IDs handed out alphabetically, or by frequency, or by whim 1 100,000 bill 4,102 subscription 88,301 a gap of 84,199 that says nothing about English
An ID is fine for filing a word and useless for comparing one. In an ID world “cat” and “kitten” are exactly as far apart as “cat” and “thermodynamics” — the gap is an artifact of the sorting.

3 · The key idea

Stop choosing the numbers. Let one rule move them.

PULL TOGETHER stop the bill cancel your subscription things that should be alike PUSH APART stop the bill reset your password things that shouldn’t be repeat over a huge pile of text, and the coordinates settle themselves
Nobody picks where a phrase lands. Run “pull together, push apart” long enough and the content draws its own map — cancellation articles drift into one neighbourhood, password resets into another.

4 · The thing the rule hides

Somebody still has to say which pairs "should" be close. Text says it for free.

A · FREE PAIRS, FROM RAW TEXT word2vec, 2013 — slide a window over a huge pile of writing … you can cancel the plan any time, or stop the bill from your … the window neighbours → “should be close” a word plucked at random → “should be far” B · CHOSEN PAIRS, FROM YOUR LOGS question and answer, paraphrases, clicks — whole passages, not words stop the bill cancel your subscription pulled together on purpose
This is the seam: you never define what “similar” means — the training pairs do. Change the pairs and you change the geometry, and nobody writes down what the geometry now believes.

5 · The shortcut

Do the expensive half once, before anyone asks.

OFFLINE, ONCE 2M help articles embed one vector each store VECTOR INDEX built to answer one question ONLINE, PER QUESTION stop the bill embed what’s nearest this pin? the nearest 8 articles paste into the model’s prompt
Comparing your query against two million articles per keystroke is impossible; comparing it against a prepared index is milliseconds. That split is the whole skeleton of retrieval-augmented generation — the model looks like the clever part, but retrieval decided what it got to see.

6 · Keep this card

The whole thing on one index card.

embedding = a learned map of meaning + similar things land near each other + “similar” = whatever the pairs said
Picture to keep: a map where you don’t choose the coordinates — the content does. Your typed question lands as a pin somewhere on it, and search is just “what’s near the pin.”

Why it exists

You want to cancel a subscription, so you open the company’s help centre and type what’s actually in your head: “stop the bill.” The right article is called “How do I cancel my subscription?” — and it comes up first. Look at those two strings. They share no words at all. “Stop” isn’t “cancel,” “bill” isn’t “subscription,” and nothing about the letters connects them.

A search box that only matched words would have returned nothing here, and you’d be emailing support. Something in that box knows the two phrases are about the same thing, and it isn’t a synonym list somebody wrote for your exact wording.

The problem underneath: a neural network can’t read the word “bill.” It multiplies numbers. So text has to become numbers somewhere. The boring way — give every word an arbitrary integer ID — is fine for indexing and destroys everything interesting. In an ID-based world “cat” and “kitten” are exactly as far apart as “cat” and “thermodynamics,” because the IDs were assigned alphabetically or by frequency or by whim.

Embeddings exist to give you numbers that keep the meaning. Two demands at once:

  1. It’s a vector of numbers, so a model or a database can do math on it.
  2. Similar things land near each other, so distance stands in for “are these alike?”

The second demand is the one that pays. Once you have it, a whole family of problems — find the doc that answers this, suggest songs like this song, cluster these tickets, flag these two reviews as near-duplicates — collapses into one operation: find the nearest points. The idea usually credited underneath all of it is distributional semantics: if “cat” and “kitten” show up around the same other words, maybe they mean similar things. Hold on to one question while you read, because it’s the whole game: who decides what “similar” means? We’ll follow “stop the bill” until that has an answer.

Why it matters now

Two different things get called embeddings, and it’s worth separating them. Every LLM has an internal embedding table that turns token IDs into vectors before the first layer — you get that for free, and it isn’t what people mean by “using embeddings.” The other kind is a separate model whose whole output is one vector per input — a document, a chunk, a query, an image, and that’s the one you deliberately reach for:

The engineer’s reason to care: an off-the-shelf embedding model plus a vector store buys you a lot of “AI features” without fine-tuning anything. That’s my opinion about cost-effectiveness rather than a measured claim — but the timeline isn’t opinion. word2vec is from 2013, well before anything anyone would call an LLM today.

The short answer

embedding = a learned function (text | image | thing) → ℝᵈ such that semantic similarity ≈ vector similarity

Picture to keep: a map where you don’t choose the coordinates — the content does. Help articles about cancelling drift into one neighbourhood, articles about password resets into another, and your typed question lands as a pin somewhere on that map. Search is “what’s near the pin.”

An embedding is a fixed-length list of floats — a few hundred to a couple of thousand is the range you’ll meet in practice, though check whatever model you’re using — produced by a model trained so that things meaning similar stuff come out pointing in similar directions. “Are these two alike?” becomes arithmetic.

How it works

Keep two questions apart — how you build the map, and how you use it. Both are easiest to see by failing your way forward.

Attempt 1: match the words. Search for “stop the bill” by looking for documents containing those words.

Why it breaks: zero overlap with “cancel your subscription.” Keyword search is excellent when the user already speaks your product’s vocabulary, and useless the moment they don’t — which is most of the time, since the people searching for how to cancel are precisely the people who never learned your terminology.

Attempt 2: write a synonym list. Map “stop” → “cancel,” “bill” → “subscription,” “invoice,” “charge.”

Why it breaks: somebody has to maintain it, forever, per language, per product, and it doesn’t compose. “Stop the bill” needs the phrase to mean cancellation; “stop” alone might be about pausing a video. You’re back to hand-writing the knowledge, which is the thing you were trying to avoid.

Attempt 3: number the words. Give every word an ID and let a model learn from those.

Why it breaks: the IDs are arbitrary. If “bill” is 4102 and “subscription” is 88,301, the gap between them is an artifact of the sorting, not a fact about English. There’s nothing to learn from; the model would have to memorize every pair separately, which is Attempt 2 with extra steps.

Fix: let the corpus assign the numbers. Instead of picking coordinates, learn them, using one rule with two halves:

Pull together representations of things that should be similar. Push apart representations of things that shouldn’t.

The trick is defining “should” without a human labelling every pair — and text supplies it free. word2vec (Mikolov et al., 2013), in its skip-gram form, slides a window over a huge corpus, treats each word’s neighbours as “should be close,” and treats randomly-drawn other words as “should be far.” Because “cancel” and “stop” appear around similar words, they end up in similar places without anyone saying so. The famous party trick — vec("king") − vec("man") + vec("woman") ≈ vec("queen") — fell out as a side effect. (How reliable that analogy really is has been argued about ever since; the durable point is that usable linear structure appeared in the geometry at all.)

Why it breaks: one vector per word. “Stop the bill” is three vectors, and averaging them throws away that this is a phrase about billing rather than a sentence containing the word “bill.” Worse, a static vector for “bill” has to serve the invoice, the legislation, and the duck.

Fix: embed the whole passage in context. Run the sentence through a transformer, which lets each word’s representation depend on its neighbours, and pool the result into a single vector. Train it on pairs that are known to mean the same thing — question/answer pairs, paraphrases, click logs — so “stop the bill” and “cancel your subscription” are pulled together explicitly. Sentence-BERT (2019) is the clearest early example of this shape; contrastive training on sentence pairs is still the common recipe behind the embedding models you’d call today, though the specific architectures and data have moved on a lot. Note what just happened to our question: “similar” now means whatever the training pairs said it means. Nobody defined it; somebody chose examples. The same recipe crosses modalities — CLIP trains an image encoder and a text encoder together on hundreds of millions of (image, caption) pairs, so pictures land near the words that describe them.

You will almost never train one of these. You pick one off the shelf and treat it as a black box.

Why it breaks: comparing your query against two million help articles, one at a time, per keystroke.

Fix: do the expensive half offline. Embed every document once, store the vectors, and at query time embed only the query and look up its neighbours — with an ANN index so “nearest 10 of 100 million” is milliseconds rather than a table scan:

# offline, once
for chunk in docs:
    store(id=chunk.id, vector=embed(chunk.text), payload=chunk)

# online, per question
q_vec  = embed("stop the bill")
hits   = vector_db.search(q_vec, top_k=8)
prompt = SYSTEM + format(hits) + user_question
answer = llm(prompt)

That’s the skeleton of RAG. Production systems layer a lot on top — keyword search blended with vector search, a reranker over the top-k, query rewriting, chunking strategy — but the retrieve-then-prompt shape is this. The LLM looks like the clever part; the embedding step is why the right eight chunks were in the prompt at all.

Where the map picture stops being safe, and the things that bite people:

You started with embedding = learned function → vectors where similar means near. What did “stop the bill” add that the definition leaves out? — you never specify what “similar” means; the training pairs do. That’s the whole reason it scales past the synonym list, and also the reason it fails in ways you didn’t choose: change the pairs and you change the geometry, and nobody wrote down what the geometry now believes.

Going deeper

A note on what I’m sure of: the high-level story — objectives, use cases, the contrastive shape, the “similar things point the same way” property — is well established. Benchmark numbers and “best embedding model right now” rankings change every few months; verify those against a current leaderboard rather than memorizing anything here.