What is an LLM?
A neural network trained to predict the next token of text — and why that simple goal scaled into something that feels like reasoning.
On this page
The picture version
Six pictures for a reader who has never heard the term. The prose below fills in the seams the pictures skip.
1 · The problem
It answers a question nobody prepared it for.
2 · The old way
One hand-built machine per task — and one is always missing.
3 · The trick
Hide the next piece of text. The answer key is the text itself.
4 · One step, up close
Chop into pieces, look back, score every possible next piece, pick one, repeat.
5 · The missing piece
A text-continuer isn’t an assistant yet.
6 · Keep this card
The whole thing on one index card.
Why it exists
You paste a red wall of stack trace into a chat box — TypeError: Cannot read properties of undefined (reading 'map'), forty lines of framework internals —
and type “what’s wrong here?” A paragraph comes back explaining that your API
returned null before the component rendered, and suggesting a guard clause.
Nobody wrote a rule for your stack trace. Nobody wrote a rule for stack traces
at all.
That’s the thing worth explaining, and I’ll keep coming back to this same stack trace for the rest of the post. Language software used to be built one task at a time — a parser and a grammar here, a hand-tuned statistical translation system there, each brittle in its own way. To handle your error message you’d have needed a parser for stack traces, a table of framework-specific diagnostics, and a maintainer to update it every release.
The hope behind LLMs is older than the technology: maybe a single model, fed enough text, could learn the structure of language by itself — without anyone sitting down to write the rules. Once it could, you wouldn’t build a separate system for translation and another for debugging. You’d just ask.
The standard account of why that took so long is that models weren’t expressive enough, data wasn’t big enough, and there wasn’t the compute to train on it — and that the transformer architecture (Vaswani et al., 2017) plus the GPU build-out moved all three at once. Whatever the exact weighting, the result is what matters: the simplest possible objective, “predict the next word,” turned out to be enough to get models that write code and explain jokes. Nobody wrote a stack-trace-explaining module. At scale, predicting the next word well seems to require a lot of the same competence.
Why it matters now
LLMs are what’s underneath the AI features people actually touch: chat assistants, coding tools, document Q&A, support bots, the suggestion bar in your editor. If you build software you will very likely end up calling one; if you don’t, you’re already using products that do.
The mechanics are worth understanding even at a high level, because the ways LLMs fail are the new bugs in the systems being shipped — and each one traces back to the mechanism below. Hallucination is largely what next-token prediction does when it has nothing to go on and no way to abstain. Prompt injection works because the model has no structural way to tell your instructions from text someone else wrote. Context limits are a memory bill. None of these are add-on defects; they’re the shape of the thing.
The short answer
LLM = neural net + "predict the next token" objective at scale
Picture to keep: an enormous autocomplete, running one fragment at a time — it never sees your whole answer, only the next piece of it. Except that your phone’s autocomplete looks at the last few words and this one looks at everything in the conversation, and a later training stage taught it that a question should be followed by an answer rather than by more question.
An LLM is a neural network trained on huge amounts of text to predict the next token (≈ word piece) given the tokens that came before. That’s the entire pretraining objective, and it’s where the general competence comes from — though as the last section shows, the assistant you actually talk to has had further training layered on top of it.
How it works
The cleanest way to see the design is to try to build it badly and watch each piece fail.
Naive attempt: train on question → answer pairs. You want a model that maps “here’s my stack trace, what’s wrong?” to an explanation. So collect millions of such pairs and train on them. This dies immediately on data: no one has labelled a billion stack traces, and any set you could assemble covers a sliver of what people ask.
Fix: predict the next token instead. Take raw text — no labels needed — hide the next chunk, and train the model to guess it. Now ordinary text is usable as training data, because the label is just “what actually came next.” (Real pretraining corpora are still heavily filtered, deduplicated and chosen; what disappears is the need for anyone to annotate them.) Explanations of stack traces exist in the wild, in blog posts and Stack Overflow answers — so continuing text well should, on this argument, mean learning to produce them too.
But what’s a “token”? Train on whole words and your vocabulary can’t hold
the long tail — TypeError, useEffect, and your variable names aren’t
words. Train on single characters and the model spends its capacity learning
spelling. The fix is subword tokenization:
chop text into frequent chunks like " the", "Type", "Error", ".", each
with an integer ID. Open models today publish vocabularies from tens of
thousands to a couple of hundred thousand entries (Llama 3, for instance,
uses 128,256), and anything rare gets spelled out of common pieces. Your
stack trace becomes a list of integers.
But which earlier tokens matter? A model that reads a fixed-size window,
or averages everything it has seen, can’t tell that the word undefined on
line 1 is what explains the word null it wants to write on line 40. The fix
is attention: at every position, each token
computes how relevant every earlier token in the context is, and pulls from
the relevant ones. In “the cat sat on the mat because it was warm”, some
head can learn to link “it” back to “mat” — that specific pairing is an
illustration of the shape, not a claim about a head anyone has labelled. A
transformer
is a stack of these attention layers alternating with feed-forward layers, a
learned per-token transformation; Llama 3.1 70B has 80 such layers. After the
last layer, the model emits a probability over the whole vocabulary for the
next token.
But a probability isn’t text. Something has to pick. That’s sampling: draw a token from the distribution (sometimes the most likely one, sometimes weighted randomly), append it to the context, and run the whole thing again. Token by token, the explanation of your stack trace appears. It’s also why, at the randomized settings most chat products ship with, the same prompt twice gives you two different paragraphs.
But now it just continues text — it doesn’t answer. A base model handed your stack trace is as likely to generate another stack trace as an explanation; on the internet, that’s often what follows. Post-training fixes this. The recipes differ by lab and keep changing, but two stages are the common shape:
- Instruction tuning — fine-tune on examples of “user asks X, assistant responds Y” so the model learns that a question should be followed by an answer, not by more question.
- Reinforcement learning from human/AI feedback (RLHF / RLAIF) — train the model to prefer responses that humans (or another model) rated higher. This is a large part of why the assistant feels helpful and mostly avoids obviously bad outputs — though it’s a preference signal, not a correctness one, which is its own set of problems.
The “intelligence” is compressed into the transformer’s weights — billions of numbers for a typical open model, and more for the largest ones — formed during pretraining on internet-scale text. Nothing in there is a rule about stack traces.
You started with LLM = neural net + "predict the next token" objective at scale. What did the walk-through add? — + subwords + attention + a sampler + a tuning pass that turns continuation into answering. Only the middle two are
the model; the sampler and the tuning are what make a text-continuation engine
feel like something you can talk to.
Famous related terms
- Transformer —
transformer ≈ stack of (attention + feed-forward) layers— the architecture that made “which earlier token matters?” learnable. - Tokenization —
tokenization = text → list of integer IDs— why the model never actually sees your words, only chunks of them. - Embeddings —
embedding = thing → vector you can do math on— how a token ID becomes something you can multiply. - Attention —
attention = each token weights every other token by learned relevance— the step that pulledundefinedon line 1 into the answer about line 40. - Context window —
context window = how many tokens the model can see at once— the hard edge of what the sampler is conditioned on. - Hallucination —
hallucination = confident output that isn't true— what next-token prediction does when it has no fact to condition on and no way to abstain. - RLHF —
RLHF = supervised fine-tune + reward model + RL loop— the step that turns a text-continuation engine into an assistant.
Going deeper
- Language Models are Few-Shot Learners (Brown et al., 2020) — the primary source for the claim this post rests on: scale a next-token predictor far enough and it starts doing tasks nobody trained it on.
- Andrej Karpathy, Let’s build GPT — if you want the loop above to stop being a metaphor, this builds a working one from an empty file in about two hours.
- Attention Is All You Need (Vaswani et al., 2017) — the rabbit hole, and the primary source for the architecture: what an attention layer precisely computes, rather than what it’s for.