10 famous AI-ML terms
The vocabulary you keep hearing on every podcast — neural network, transformer, RLHF, RAG — compressed to one line each, then unpacked.
On this page
- The picture version
- Why it exists
- Why it matters now
- The short answer
- How it works
- 1. Tokenization
- 2. Embedding
- 3. Neural network
- 4. Transformer
- 5. Attention
- 6. LLM (Large Language Model)
- 7. RLHF (Reinforcement Learning from Human Feedback)
- 8. Hallucination
- 9. RAG (Retrieval-Augmented Generation)
- 10. Agent
- Famous related terms
- Going deeper
The picture version
Six pictures for a reader who has never met any of these words. They follow one question — can I expense a train ticket? — from your keyboard to the answer.
1 · The problem
Two seconds pass. Ten famous words live in that gap.
2 · The first two jobs
Your words stop being words, then the numbers get meaning.
3 · The machine
Stack the same block many times, guess one token.
4 · Where it breaks
A next-token guesser disappoints you in two ways.
5 · The fixes stack up
Nothing you actually talk to is just the base model.
6 · Keep this card
Ten names to memorize, ten reasons you can rebuild.
Why it exists
You type a question into a work chatbot — “can I expense a train ticket to a client meeting?” — and two seconds later you have an answer with a link to the travel policy. That single request is the running example for this whole post, because every one of the ten famous words below names a job that happened between your Enter key and that answer.
Which is the problem this post fixes. You’ve collected these words — token, transformer, embedding, RLHF, hallucination, agent — from podcasts and press releases, where they get used as if everyone already knows them. They sound technical because they are, but most of them compress to a single sentence once someone draws the picture. Ten pictures, in the order your question actually meets them.
Why it matters now
These words now describe things you use and things you’re asked to decide about. The chat assistant in your company’s help desk, the “summarize this thread” button in your email, the autocomplete in your editor, the tool your team is being asked to buy this quarter — the differences between them are mostly differences in these ten words. We use a transformer tells you almost nothing, because nearly every current system does; we use RAG over your documents tells you where the answers come from, what happens when the document is missing, and who has to keep the index fresh.
I picked these ten because they’re the ones non-specialists hit most often when trying to discuss modern AI. Other strong candidates — fine-tuning, chain-of-thought, diffusion, MoE, quantization — are in Famous related terms at the bottom.
The short answer
AI-ML jargon = math you can re-derive + branding you have to memorize
Picture to keep: a relay race from your keystroke to the answer on screen, where each of these words names one runner and the baton is your question, in progressively less recognizable form.
Most of the famous terms are an old idea — matrix multiplication, gradient descent, probability — wearing a new name plus one specific recipe. The recipe is the part worth knowing; the name is just what you have to memorize to read the room.
How it works
Follow the train-ticket question down the track.
1. Tokenization
tokenization = text → list of integer IDs the model actually sees
Your question never reaches the model as letters. The obvious schemes both fail:
one entry per word and the model is helpless at any word it never saw; one entry
per character and every sentence becomes punishingly long. So text is chopped
into tokens from a fixed vocabulary — typically tens of thousands of entries —
and each becomes an integer. A token might be " the", "un", or "derstand"; “expense” might be
one token or two depending on the vocabulary. This is why models miscount the
letters in a word, why some languages cost more per request than others, and why
“context length” is counted in tokens rather than words. See
tokenization.
2. Embedding
embedding = a thing → a vector you can do math on
Those integer IDs are arbitrary — ID 4102 tells you nothing about ID 4103 — so each one is looked up in a table of learned vectors. Similar things end up pointing in similar directions, which is what lets “expense” and “reimbursement” be treated as related without anyone writing a synonym list. The same idea, applied to whole documents, is what will later find the travel policy. See embeddings.
3. Neural network
neural network = stack of (linear layer + nonlinearity) + gradient descent
Vectors are still just numbers; something has to turn them into a prediction. Each “neuron” multiplies its inputs by weights, adds a bias, and then bends the result with a simple curve — a nonlinearity. Stack enough of those and the whole thing can fit patterns nobody knows how to write down by hand. The compression line above deliberately packs two things together — the shape (layers and nonlinearities) and the training method (gradient descent: nudge every weight, one mini-batch at a time, in whatever direction reduces the error). They’re separable in principle; in practice you never meet one without the other. See neural network.
4. Transformer
transformer = stack of (self-attention + feed-forward) blocks + positional info
Which kind of neural network handles your sentence matters. Older recurrent networks marched through a sentence one word at a time; a transformer computes all positions in one parallel operation, each able to look at the others. (In the decoder-only models behind chat assistants, “the others” means only earlier positions — a mask blocks the future, or predicting the next token would be trivial.) Parallelism is the point: it’s what lets training saturate a GPU cluster, and that efficiency is a large part of why this architecture, rather than a serial one, scaled the way it did. See transformer.
5. Attention
attention = each position scores the positions it can see, then blends them
The operation inside a transformer block that does the looking. Each position asks “which other positions matter to me right now?” and pulls in a weighted blend of their information. The scores aren’t a fixed lookup table — they’re computed from the current activations every time, at every layer, which is why it can refer to the train ticket in one sentence and to the client meeting in the next. What’s learned is how to compute the scores, not the scores themselves. See attention.
6. LLM (Large Language Model)
LLM = neural net + "predict the next token" objective at scale
Put those blocks together, train on an enormous amount of text with the single job of guessing the next token, and you have the base model. The surprising part is how much general capability falls out of that one narrow objective. The part this simplification hides: nothing you actually talk to is just a base model. Instruction tuning, preference tuning, retrieval, and tool scaffolding all sit on top, and the next few entries are exactly those layers. See LLM.
7. RLHF (Reinforcement Learning from Human Feedback)
RLHF = supervised fine-tune + reward model + RL loop
Here’s a thing worth noticing: a pure next-token predictor, given your question, might reasonably continue with another question — that’s what a page of FAQs looks like. Post-training is what makes it answer instead. First, supervised fine-tuning on demonstrations of good answers; that alone does much of the “respond rather than continue” work. Then the reinforcement-learning half: humans rank pairs of responses, a reward model learns to predict those rankings, and the model is optimized against it — which is more about which answer, and how it’s phrased, than about answering at all. That’s why the compression line names three moves rather than one. See why RLHF exists.
8. Hallucination
hallucination = confident output that isn't true
Assume your company’s internal travel policy was never in the training data — usually a safe bet for an internal document. Ask anyway and the model still produces fluent text, possibly citing “Section 4.2 of the Employee Travel Guidelines,” which does not exist. There’s no internal “do I actually know this?” gate; next-token prediction always yields something, and near the edge of what the model saw, that something is often invented. See hallucination.
9. RAG (Retrieval-Augmented Generation)
RAG = retrieve relevant docs + stuff them into the prompt + generate
So don’t ask it to remember. Before answering, search your actual policy documents — often by embedding similarity, from step 2, sometimes blended with plain keyword search — and paste the most relevant chunks into the prompt alongside the question. Now there’s real evidence in front of the model, which is why the answer could come with a link. Note what it doesn’t do: the model is still using everything it learned in training — RAG adds evidence rather than switching memory off. That’s why it reduces hallucination without abolishing it, and why a wrong retrieval produces a confident wrong answer. See RAG.
10. Agent
agent = model + harness (tools + loop + memory + control flow)
One retrieval and one answer covered the train ticket. Now ask it to file the expense: check the policy, find the receipt, fill the form, flag the exception. That needs a loop — call a tool, read the result, decide what’s next — wrapped around the model, with rules about what it’s allowed to do unattended. The buzz is about the model; most of the engineering is that wrapper. See agent harness.
You started with AI-ML jargon = math you can re-derive + branding you have to memorize. Having run one question down the track, which half turns out to be
bigger? — the re-derivable half, and the ordering is why. Read as a chain,
most of these terms are a fix for a problem the previous one left behind: IDs are
arbitrary → embed them; recurrence is serial → attend in parallel; a text
continuer won’t answer → post-train it; it doesn’t know your documents → retrieve
them; one answer isn’t a task → wrap it in a loop. That chain is a lens I’ve
chosen for teaching, not the literal order of history — several of these were
developed in parallel by different groups. But it means you only have to memorize
ten names; the ten reasons you can rebuild from the failures.
Famous related terms
- Fine-tuning —
fine-tuning = continue training a pretrained model on your own data— cheap because pretraining did the hard part; see why fine-tuning is cheap. - Chain-of-thought —
CoT = let the model write its reasoning before its answer— often improves accuracy on multi-step problems, and costs tokens; see chain-of-thought. - Diffusion model —
diffusion = learn to denoise + run the denoiser repeatedly— the generative recipe behind most current image and video systems; see why diffusion models exist. - Mixture of Experts (MoE) —
MoE = many expert FFNs + a router that picks a few per token— how recent frontier models grow capacity without paying full per-token compute; see why mixture of experts exists. - Quantization —
quantization = store weights (and often activations) in fewer bits instead of 16/32— makes big models fit on smaller hardware; see why quantization works. - Context window —
context window = how many tokens the model can see at once— bigger isn’t automatically better; see why lost in the middle.
Going deeper
- Vaswani et al., Attention Is All You Need (2017) — answers “what is a transformer block, precisely?” without any of the hand-waving above, in about eight pages.
- Andrej Karpathy, Let’s build GPT: from scratch, in code, spelled out — answers “could I build one myself?”; terms 1 through 6 stop being words once you’ve typed them as code.
- Ouyang et al., Training language models to follow instructions with human feedback (2022) — answers “what exactly did they do to make it behave like an assistant?” This is the InstructGPT paper behind the supervised-fine-tune + reward-model + PPO recipe. The RLHF idea itself is older: Christiano et al. (2017) and Stiennon et al. (2020).