Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

What is an LLM?

A neural network trained to predict the next token of text — and why that simple goal scaled into something that feels like reasoning.

AI & ML intro Apr 29, 2026 · updated Aug 25, 2026 · 9 min read

On this page

The picture version

Six pictures for a reader who has never heard the term. The prose below fills in the seams the pictures skip.

1 · The problem

It answers a question nobody prepared it for.

TypeError: Cannot read properties of undefined (reading ‘map’) at ProductList (list.jsx:41:18) at renderWithHooks (react-dom:16305) at mountIndeterminateComponent … 37 more lines something you pasted in without explaining it ask Your list arrived as nothing at all, because the request hadn’t finished when the page tried to draw it. Check for that case before drawing the list. a plain-English answer, in seconds No one ever wrote a rule for this error message.
The reply isn’t looked up anywhere. Nothing in the system is a rule about error messages — and yet an explanation comes back. That is the thing worth explaining.

2 · The old way

One hand-built machine per task — and one is always missing.

translation spell check grammar rules hand-written hand-written hand-written stack-trace explainer nobody built this one — and who’d maintain it? Each one brittle in its own way, each needing its own experts. ONE TASK · ONE BUILD · NEVER ENOUGH
Building a system per task means every new question needs a new build, a labelled dataset, and someone to keep it current. The long tail of things people actually ask never gets covered.

3 · The trick

Hide the next piece of text. The answer key is the text itself.

This error usually means the data hadn’t arrived yet, so the safest fix is to covered up guess this check for it before drawing the right answer was already sitting there — it’s just what came next and there is a lot of ordinary text no one has to mark up anything — every passage grades its own guess
Instead of collecting question-and-answer pairs somebody wrote by hand, cover the next chunk of ordinary text and train the system to guess it. The label is free, so the training material is enormous — and explanations of errors already live inside it.

4 · One step, up close

Chop into pieces, look back, score every possible next piece, pick one, repeat.

Type Error : Can not undefin ed … your text, chopped into frequent chunks (whole words, or parts of them) ? looks back at whichever earlier chunks matter scores every chunk it knows “ is” “ means” “ usually” “ the” … how likely each one is to come next pick one “ usually” stick it on the end and run the whole thing again — one chunk at a time, until it stops The picking step is outside the model, and usually a bit random — which is why the same question twice gives two different paragraphs.
The model never plans the paragraph. It scores what could come next, something picks one, and the loop runs again with the answer growing on the end. The explanation of your error is built one chunk at a time, from the left.

5 · The missing piece

A text-continuer isn’t an assistant yet.

straight out of pretraining TypeError: Cannot read properties… TypeError: Cannot read properties of undefined (reading ‘length’) on the internet, an error message is often followed by another one after post-training TypeError: Cannot read properties… Your list arrived as nothing at all… here is what to check. shown thousands of request-and-good-answer pairs, and steered toward answers people preferred Same machine, same trick. What changed is what it treats a question as.
The general competence comes from predicting text; the habit of answering comes from a later training pass. That later pass rewards responses people rated higher — which is a preference signal, not a correctness one.

6 · Keep this card

The whole thing on one index card.

LLM = a big network of learned numbers + one goal: predict the next chunk + scale — enough text and enough numbers — the picker and the tuning pass are what make it talk back
Picture to keep: an enormous autocomplete, running one fragment at a time — it never sees your whole answer, only the next piece of it.

Why it exists

You paste a red wall of stack trace into a chat box — TypeError: Cannot read properties of undefined (reading 'map'), forty lines of framework internals — and type “what’s wrong here?” A paragraph comes back explaining that your API returned null before the component rendered, and suggesting a guard clause. Nobody wrote a rule for your stack trace. Nobody wrote a rule for stack traces at all.

That’s the thing worth explaining, and I’ll keep coming back to this same stack trace for the rest of the post. Language software used to be built one task at a time — a parser and a grammar here, a hand-tuned statistical translation system there, each brittle in its own way. To handle your error message you’d have needed a parser for stack traces, a table of framework-specific diagnostics, and a maintainer to update it every release.

The hope behind LLMs is older than the technology: maybe a single model, fed enough text, could learn the structure of language by itself — without anyone sitting down to write the rules. Once it could, you wouldn’t build a separate system for translation and another for debugging. You’d just ask.

The standard account of why that took so long is that models weren’t expressive enough, data wasn’t big enough, and there wasn’t the compute to train on it — and that the transformer architecture (Vaswani et al., 2017) plus the GPU build-out moved all three at once. Whatever the exact weighting, the result is what matters: the simplest possible objective, “predict the next word,” turned out to be enough to get models that write code and explain jokes. Nobody wrote a stack-trace-explaining module. At scale, predicting the next word well seems to require a lot of the same competence.

Why it matters now

LLMs are what’s underneath the AI features people actually touch: chat assistants, coding tools, document Q&A, support bots, the suggestion bar in your editor. If you build software you will very likely end up calling one; if you don’t, you’re already using products that do.

The mechanics are worth understanding even at a high level, because the ways LLMs fail are the new bugs in the systems being shipped — and each one traces back to the mechanism below. Hallucination is largely what next-token prediction does when it has nothing to go on and no way to abstain. Prompt injection works because the model has no structural way to tell your instructions from text someone else wrote. Context limits are a memory bill. None of these are add-on defects; they’re the shape of the thing.

The short answer

LLM = neural net + "predict the next token" objective at scale

Picture to keep: an enormous autocomplete, running one fragment at a time — it never sees your whole answer, only the next piece of it. Except that your phone’s autocomplete looks at the last few words and this one looks at everything in the conversation, and a later training stage taught it that a question should be followed by an answer rather than by more question.

An LLM is a neural network trained on huge amounts of text to predict the next token (≈ word piece) given the tokens that came before. That’s the entire pretraining objective, and it’s where the general competence comes from — though as the last section shows, the assistant you actually talk to has had further training layered on top of it.

How it works

The cleanest way to see the design is to try to build it badly and watch each piece fail.

Naive attempt: train on question → answer pairs. You want a model that maps “here’s my stack trace, what’s wrong?” to an explanation. So collect millions of such pairs and train on them. This dies immediately on data: no one has labelled a billion stack traces, and any set you could assemble covers a sliver of what people ask.

Fix: predict the next token instead. Take raw text — no labels needed — hide the next chunk, and train the model to guess it. Now ordinary text is usable as training data, because the label is just “what actually came next.” (Real pretraining corpora are still heavily filtered, deduplicated and chosen; what disappears is the need for anyone to annotate them.) Explanations of stack traces exist in the wild, in blog posts and Stack Overflow answers — so continuing text well should, on this argument, mean learning to produce them too.

But what’s a “token”? Train on whole words and your vocabulary can’t hold the long tail — TypeError, useEffect, and your variable names aren’t words. Train on single characters and the model spends its capacity learning spelling. The fix is subword tokenization: chop text into frequent chunks like " the", "Type", "Error", ".", each with an integer ID. Open models today publish vocabularies from tens of thousands to a couple of hundred thousand entries (Llama 3, for instance, uses 128,256), and anything rare gets spelled out of common pieces. Your stack trace becomes a list of integers.

But which earlier tokens matter? A model that reads a fixed-size window, or averages everything it has seen, can’t tell that the word undefined on line 1 is what explains the word null it wants to write on line 40. The fix is attention: at every position, each token computes how relevant every earlier token in the context is, and pulls from the relevant ones. In “the cat sat on the mat because it was warm”, some head can learn to link “it” back to “mat” — that specific pairing is an illustration of the shape, not a claim about a head anyone has labelled. A transformer is a stack of these attention layers alternating with feed-forward layers, a learned per-token transformation; Llama 3.1 70B has 80 such layers. After the last layer, the model emits a probability over the whole vocabulary for the next token.

But a probability isn’t text. Something has to pick. That’s sampling: draw a token from the distribution (sometimes the most likely one, sometimes weighted randomly), append it to the context, and run the whole thing again. Token by token, the explanation of your stack trace appears. It’s also why, at the randomized settings most chat products ship with, the same prompt twice gives you two different paragraphs.

But now it just continues text — it doesn’t answer. A base model handed your stack trace is as likely to generate another stack trace as an explanation; on the internet, that’s often what follows. Post-training fixes this. The recipes differ by lab and keep changing, but two stages are the common shape:

The “intelligence” is compressed into the transformer’s weights — billions of numbers for a typical open model, and more for the largest ones — formed during pretraining on internet-scale text. Nothing in there is a rule about stack traces.

You started with LLM = neural net + "predict the next token" objective at scale. What did the walk-through add? — + subwords + attention + a sampler + a tuning pass that turns continuation into answering. Only the middle two are the model; the sampler and the tuning are what make a text-continuation engine feel like something you can talk to.

Going deeper