Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

10 famous AI-ML terms

The vocabulary you keep hearing on every podcast — neural network, transformer, RLHF, RAG — compressed to one line each, then unpacked.

AI & ML intro Apr 30, 2026 · updated Aug 24, 2026 · 10 min read

On this page

The picture version

Six pictures for a reader who has never met any of these words. They follow one question — can I expense a train ticket? — from your keyboard to the answer.

1 · The problem

Two seconds pass. Ten famous words live in that gap.

you can I expense a train ticket? an answer + a link to the travel policy two seconds later ? ten famous words name what happens in here
Between your Enter key and that answer, ten separate jobs run. Each of the ten famous words names one of them — and they happen in a fixed order.

2 · The first two jobs

Your words stop being words, then the numbers get meaning.

can I expense a train ticket? chopped up can I ex pense a train ticket ? one possible split — vocabularies differ 4102 4103 IDs are arbitrary 4102 tells you nothing about 4103 look up expense reimbursement learned vectors
Tokenization turns the sentence into integer IDs; embedding turns each ID into a vector. That second step is why expense and reimbursement can point the same way without anyone writing a synonym list.

3 · The machine

Stack the same block many times, guess one token.

the next token one guess self-attention + feed-forward self-attention + feed-forward self-attention + feed-forward × many layers looks back at earlier positions can I ex pense a train one vector per token
A neural network is a stack of layers; a transformer is a stack of this block, computed for all positions at once. Attention is the looking-back arrows. Train the whole thing to guess the next token, at scale, and you have an LLM.

4 · Where it breaks

A next-token guesser disappoints you in two ways.

1 · it continues can I expense a train ticket? continues and what about hotels? and what about taxis? a page of FAQs, not an answer fixed by post-training (RLHF) 2 · it invents what is our travel policy? never saw it Section 4.2 of the Employee Travel Guidelines fluent, confident, does not exist fixed by retrieval (RAG) nothing inside asks “do I actually know this?”
Neither failure is a bug in the math. Next-token prediction always yields something, so the raw model continues where you wanted an answer, and invents where you wanted a citation.

5 · The fixes stack up

Nothing you actually talk to is just the base model.

everything bolted on top agent harness RAG (retrieval) RLHF + instruction tuning base LLM tools + loop + memory + control flow puts your documents in the prompt makes it answer, not continue guesses the next token most of the engineering is here adds evidence, doesn’t switch memory off fixes: it continued instead of answering the buzz is about this box
The base model only guesses the next token. Post-training makes it answer, retrieval hands it your documents, and the harness turns one answer into a task — each layer is a repair for the layer below it.

6 · Keep this card

Ten names to memorize, ten reasons you can rebuild.

AI-ML jargon = math you can re-derive + branding you have to memorize and the first half is the bigger one
Picture to keep: a relay race from your keystroke to the answer on screen, where each famous word names one runner and the baton is your question in progressively less recognizable form. Memorize the ten names; rebuild the ten reasons from the failures.

Why it exists

You type a question into a work chatbot — “can I expense a train ticket to a client meeting?” — and two seconds later you have an answer with a link to the travel policy. That single request is the running example for this whole post, because every one of the ten famous words below names a job that happened between your Enter key and that answer.

Which is the problem this post fixes. You’ve collected these words — token, transformer, embedding, RLHF, hallucination, agent — from podcasts and press releases, where they get used as if everyone already knows them. They sound technical because they are, but most of them compress to a single sentence once someone draws the picture. Ten pictures, in the order your question actually meets them.

Why it matters now

These words now describe things you use and things you’re asked to decide about. The chat assistant in your company’s help desk, the “summarize this thread” button in your email, the autocomplete in your editor, the tool your team is being asked to buy this quarter — the differences between them are mostly differences in these ten words. We use a transformer tells you almost nothing, because nearly every current system does; we use RAG over your documents tells you where the answers come from, what happens when the document is missing, and who has to keep the index fresh.

I picked these ten because they’re the ones non-specialists hit most often when trying to discuss modern AI. Other strong candidates — fine-tuning, chain-of-thought, diffusion, MoE, quantization — are in Famous related terms at the bottom.

The short answer

AI-ML jargon = math you can re-derive + branding you have to memorize

Picture to keep: a relay race from your keystroke to the answer on screen, where each of these words names one runner and the baton is your question, in progressively less recognizable form.

Most of the famous terms are an old idea — matrix multiplication, gradient descent, probability — wearing a new name plus one specific recipe. The recipe is the part worth knowing; the name is just what you have to memorize to read the room.

How it works

Follow the train-ticket question down the track.

1. Tokenization

tokenization = text → list of integer IDs the model actually sees

Your question never reaches the model as letters. The obvious schemes both fail: one entry per word and the model is helpless at any word it never saw; one entry per character and every sentence becomes punishingly long. So text is chopped into tokens from a fixed vocabulary — typically tens of thousands of entries — and each becomes an integer. A token might be " the", "un", or "derstand"; “expense” might be one token or two depending on the vocabulary. This is why models miscount the letters in a word, why some languages cost more per request than others, and why “context length” is counted in tokens rather than words. See tokenization.

2. Embedding

embedding = a thing → a vector you can do math on

Those integer IDs are arbitrary — ID 4102 tells you nothing about ID 4103 — so each one is looked up in a table of learned vectors. Similar things end up pointing in similar directions, which is what lets “expense” and “reimbursement” be treated as related without anyone writing a synonym list. The same idea, applied to whole documents, is what will later find the travel policy. See embeddings.

3. Neural network

neural network = stack of (linear layer + nonlinearity) + gradient descent

Vectors are still just numbers; something has to turn them into a prediction. Each “neuron” multiplies its inputs by weights, adds a bias, and then bends the result with a simple curve — a nonlinearity. Stack enough of those and the whole thing can fit patterns nobody knows how to write down by hand. The compression line above deliberately packs two things together — the shape (layers and nonlinearities) and the training method (gradient descent: nudge every weight, one mini-batch at a time, in whatever direction reduces the error). They’re separable in principle; in practice you never meet one without the other. See neural network.

4. Transformer

transformer = stack of (self-attention + feed-forward) blocks + positional info

Which kind of neural network handles your sentence matters. Older recurrent networks marched through a sentence one word at a time; a transformer computes all positions in one parallel operation, each able to look at the others. (In the decoder-only models behind chat assistants, “the others” means only earlier positions — a mask blocks the future, or predicting the next token would be trivial.) Parallelism is the point: it’s what lets training saturate a GPU cluster, and that efficiency is a large part of why this architecture, rather than a serial one, scaled the way it did. See transformer.

5. Attention

attention = each position scores the positions it can see, then blends them

The operation inside a transformer block that does the looking. Each position asks “which other positions matter to me right now?” and pulls in a weighted blend of their information. The scores aren’t a fixed lookup table — they’re computed from the current activations every time, at every layer, which is why it can refer to the train ticket in one sentence and to the client meeting in the next. What’s learned is how to compute the scores, not the scores themselves. See attention.

6. LLM (Large Language Model)

LLM = neural net + "predict the next token" objective at scale

Put those blocks together, train on an enormous amount of text with the single job of guessing the next token, and you have the base model. The surprising part is how much general capability falls out of that one narrow objective. The part this simplification hides: nothing you actually talk to is just a base model. Instruction tuning, preference tuning, retrieval, and tool scaffolding all sit on top, and the next few entries are exactly those layers. See LLM.

7. RLHF (Reinforcement Learning from Human Feedback)

RLHF = supervised fine-tune + reward model + RL loop

Here’s a thing worth noticing: a pure next-token predictor, given your question, might reasonably continue with another question — that’s what a page of FAQs looks like. Post-training is what makes it answer instead. First, supervised fine-tuning on demonstrations of good answers; that alone does much of the “respond rather than continue” work. Then the reinforcement-learning half: humans rank pairs of responses, a reward model learns to predict those rankings, and the model is optimized against it — which is more about which answer, and how it’s phrased, than about answering at all. That’s why the compression line names three moves rather than one. See why RLHF exists.

8. Hallucination

hallucination = confident output that isn't true

Assume your company’s internal travel policy was never in the training data — usually a safe bet for an internal document. Ask anyway and the model still produces fluent text, possibly citing “Section 4.2 of the Employee Travel Guidelines,” which does not exist. There’s no internal “do I actually know this?” gate; next-token prediction always yields something, and near the edge of what the model saw, that something is often invented. See hallucination.

9. RAG (Retrieval-Augmented Generation)

RAG = retrieve relevant docs + stuff them into the prompt + generate

So don’t ask it to remember. Before answering, search your actual policy documents — often by embedding similarity, from step 2, sometimes blended with plain keyword search — and paste the most relevant chunks into the prompt alongside the question. Now there’s real evidence in front of the model, which is why the answer could come with a link. Note what it doesn’t do: the model is still using everything it learned in training — RAG adds evidence rather than switching memory off. That’s why it reduces hallucination without abolishing it, and why a wrong retrieval produces a confident wrong answer. See RAG.

10. Agent

agent = model + harness (tools + loop + memory + control flow)

One retrieval and one answer covered the train ticket. Now ask it to file the expense: check the policy, find the receipt, fill the form, flag the exception. That needs a loop — call a tool, read the result, decide what’s next — wrapped around the model, with rules about what it’s allowed to do unattended. The buzz is about the model; most of the engineering is that wrapper. See agent harness.

You started with AI-ML jargon = math you can re-derive + branding you have to memorize. Having run one question down the track, which half turns out to be bigger? — the re-derivable half, and the ordering is why. Read as a chain, most of these terms are a fix for a problem the previous one left behind: IDs are arbitrary → embed them; recurrence is serial → attend in parallel; a text continuer won’t answer → post-train it; it doesn’t know your documents → retrieve them; one answer isn’t a task → wrap it in a loop. That chain is a lens I’ve chosen for teaching, not the literal order of history — several of these were developed in parallel by different groups. But it means you only have to memorize ten names; the ten reasons you can rebuild from the failures.

Going deeper