Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

What is a transformer?

The neural network architecture behind essentially every modern LLM — and the one big idea that made it work: drop recurrence, let every token look at every other token directly.

AI & ML intro Apr 30, 2026 · updated Aug 25, 2026 · 9 min read

On this page

The picture version

Six pictures for a reader who has never seen the word outside a headline. The prose below fills in the seams the pictures skip.

1 · The problem

“Shorten the second paragraph.” It knows which one.

the block of text you pasted ¶ 2 is in here … ten messages of other conversation since … “shorten the second paragraph” has to reach back past all of it — and does The generation of models before this one was bad at exactly that.
Reaching backward across a long conversation is the running example for this post. The transformer exists because of one change that made that reach cheap, and everything else in the architecture is repair work around that change.

2 · The old way

A conveyor belt: one token at a time, state handed forward.

step 1 step 2 step 3 … step 99 step 100 each step hands a hidden state to the next problem 1: it is a chain step 100 cannot start until step 99 has finished a GPU has thousands of cores and nothing to give all of them at once problem 2: the far end fades how much of step 1 survives early late Your paragraph had to survive every intervening step without being overwritten.
Both problems come from the same design choice: information travels along the sequence. Buying more hardware did not help, because the bottleneck was the dependency, not the throughput. (The fading curve is a sketch of the effect, not measured data.)

3 · The swap

Throw the belt away. Let every token look at every other token.

¶2 … … … … “that” one hop, however far back it sits and every pair, all at once what you gain nothing to get squeezed through — and the whole thing runs in parallel on a GPU what you pay every token sees every token means N² pairs That is the whole reason the transformer exists. The rest is repair work.
Self-attention lets each position build a weighted sum of every other position, with the weights learned and content-dependent. The cost grows quadratically with sequence length — which is most of why long context is expensive.

4 · Three things the swap broke

Each remaining part is a repair for the previous part’s damage.

breaks fix nothing knows where it is “the dog bit the man” = “the man bit the dog” inject position into the input every token arrives already stamped with where in the sequence it sits gathering is not thinking attention returns a weighted average of things already in the sequence — nothing new a small network per position gather sideways, then think privately. that pair is the transformer block. a deep stack will not train gradients pushed back through dozens of blocks vanish or explode; scales drift between blocks residuals + normalisation a straight path for the gradient top to bottom, and the scales held in range
Read top to bottom and the architecture reconstructs itself: attention is the idea, and the other three parts are the bill it came with. The last row is the least glamorous and most of the reason “just stack more blocks” became a viable strategy at all.

5 · The shape it all adds up to

One lane per token, running straight up.

¶2 … … “that” one lane per token — the residual stream attention — lanes read sideways from each other feed-forward — each lane thinks privately attention feed-forward … the same block, stacked N times … every block adds into the lane the next token at the top, the last lane is projected into a score for every word in the vocabulary
Nothing is passed along the sequence; everything is written into a lane and read across. Where the picture breaks: the lanes are not sealed pipes carrying one token’s meaning — after the first floor each lane holds a blend of whatever it read from the others. Sample the top, append, run the whole thing again: that is generation.

6 · Keep this card

The whole thing on one index card.

transformer = a stack of (attention + feed-forward) blocks + position baked into the input + residuals and normalisation holding it up ∴ a distant token is one hop away, not hundreds
Picture to keep: one lane per token running straight up through the stack, and at every floor the lanes reach sideways and copy from each other before each goes off and thinks privately. Your “second paragraph” is reachable because attention connects any two positions directly, no matter how far apart they sit.

Why it exists

Ten messages into a chat, you paste a long block of text, discuss it for a while, and then type “shorten the second paragraph.” The model knows which paragraph. It reaches back past everything you’ve said since and picks the right one. Hold onto that — reaching backward across a long conversation is the running example for this post, and it’s precisely the thing the previous generation of sequence models was bad at.

Before transformers, the default way to handle a sequence — a sentence, a time series, an audio clip — was a RNN or its better-behaved cousin, the LSTM. You fed the model one token, it updated a hidden state, you fed it the next token, it updated again. Sentence as conveyor belt.

That design had two problems that compounded each other.

The first was parallelism. Step t needs the hidden state from t−1, which needs t−2, and so on. You couldn’t really use a GPU the way GPUs want to be used — thousands of cores at once — because the work was a chain. Buying more hardware didn’t help much, because the bottleneck was the dependency, not the throughput.

The second was long-range dependencies. By the time information from the start of a passage reached the end, it had been squeezed through hundreds of tiny update steps and mostly washed out. LSTMs helped but didn’t solve it. Connecting your “second paragraph” back to a block of text ten messages ago meant fighting the architecture — the information had to survive every intervening step of the conveyor belt without being overwritten.

The 2017 paper Attention Is All You Need (Vaswani et al.) made a sharper move than people expected: throw recurrence out entirely. Instead of carrying state forward step by step, let every token look directly at every other token, in parallel, in a single operation called self-attention. No conveyor belt: a distant token is one hop away instead of hundreds, so there’s nothing for it to get squeezed through. And the whole operation is embarrassingly parallel on a GPU.

That is the whole reason the transformer exists. The rest is engineering around that one swap.

Why it matters now

Once people had an architecture that scaled cleanly with compute, scale itself became the lever — and the transformer turned out to be ridiculously general. The same block design sits underneath the current LLM families (GPT, Claude, Gemini, Llama), the ViT line of vision models, and OpenAI’s Whisper for speech recognition. Attention-based blocks are also central to protein structure prediction — AlphaFold is built around them, though the full system is a good deal more than a stack of transformer blocks. If you’re trying to understand a current foundation model above the API layer, “transformer” is very likely the shape you’re looking at.

The short answer

transformer = stack of (self-attention + feed-forward) blocks, with positional info baked into the input

Picture to keep: one lane per token running straight up through the stack, and at every floor the lanes are allowed to reach sideways and copy from each other before each one goes off and thinks privately. Nothing is passed along the sequence; everything is written into a lane and read across. Where the picture breaks: the lanes aren’t sealed pipes carrying one token’s meaning, since after the first floor each lane holds a blend of whatever it read from the others.

A transformer turns a sequence of tokens into a sequence of contextual vectors by repeatedly mixing information across positions (self-attention) and then transforming each position individually (feed-forward). Stack that block enough times, train at scale, and you get the substrate behind today’s foundation models.

How it works

Take the one swap — no recurrence, every token looks at every other token — and follow what it forces. Each piece of the architecture is the repair for the previous piece’s damage.

Start: let every token look at every other token. Text is split into tokens (see tokenization) and each token ID looks up a vector in a learned table — its embedding, so a message becomes a matrix. Self-attention then lets each position build a weighted sum of every other position, weights learned and content-dependent. “Second paragraph” can read directly from the paragraph, in one step, however far back it sits. No conveyor belt to survive.

But now the model is order-blind. Shuffle the input vectors and self-attention gives you the same set of outputs, just shuffled to match — no position has any way to know it came first. “The dog bit the man” and “the man bit the dog” are the same bag of tokens, which is fatal for language. Fix: inject position into the input. The original paper used fixed sinusoidal positional encodings; many current decoder-only LLMs use rotary variants instead. See why positional encodings exist and why RoPE replaced sinusoidal.

But gathering isn’t thinking. What attention hands back is a weighted average of (linearly transformed) things already in the sequence — it’s very good at moving information sideways and poor at turning it into something new. Fix: after each attention step, run each position through a small MLP, independently. Gather, then think. That pair — attention plus feed-forward — is the transformer block, and you stack it N times so information can make several hops.

But a stack that deep is hard to train. Gradients pushed back through dozens of blocks tend to vanish or explode, and each block’s output can drift to a scale the next one wasn’t expecting. Fix: wrap every sublayer in a residual connection so the gradient has a straight path from top to bottom, and a layer norm to keep the scales in range. These two are not decoration — they’re most of the reason “just stack more blocks” became a viable strategy at all.

The residual connection is also what makes the picture-to-keep literal. Each token’s lane is the residual stream: a vector that travels straight up the stack, with every block adding into it. Attention reads and writes back; the FFN reads and writes back. The stream is the working memory.

Finally, turn a lane back into a word. At the top, each token’s vector is projected into a distribution over the vocabulary. For an LLM, the distribution over the last position is the next-token prediction. Sample, append, run the whole thing again — that’s generation.

I’m being deliberately loose about the math (queries, keys, values, softmax, head counts, hidden sizes). The shape is re-derivable; the hyperparameters change with every new model and aren’t worth memorizing.

There is a bill for all this, and it’s worth knowing you’re paying it: letting every token see every other token means N² pairs, so cost grows quadratically with sequence length — which is a post of its own, and most of why long context is expensive.

You started with transformer = stack of (attention + feed-forward) blocks + positional info. What did this post add? — + residuals and normalization, the least glamorous part of the list and most of the reason the word stack is allowed to mean sixty layers instead of six. Meanwhile the hook resolves in the first line of the mechanism: your “second paragraph” is reachable because attention connects any two positions directly, no matter how far apart. The depth on top of that buys refinement, not reach — that’s my read of what the extra layers are doing, not something the architecture states.

Going deeper

A note on what I’m sure of: the high-level shape — embeddings, positional info, stacked attention + FFN blocks with residuals and norm, final projection — is the same across essentially every modern transformer. The details (head counts, hidden sizes, exact norm placement, positional scheme, FFN activation) vary model to model and aren’t pinned down here.