What is a transformer?
The neural network architecture behind essentially every modern LLM — and the one big idea that made it work: drop recurrence, let every token look at every other token directly.
On this page
The picture version
Six pictures for a reader who has never seen the word outside a headline. The prose below fills in the seams the pictures skip.
1 · The problem
“Shorten the second paragraph.” It knows which one.
2 · The old way
A conveyor belt: one token at a time, state handed forward.
3 · The swap
Throw the belt away. Let every token look at every other token.
4 · Three things the swap broke
Each remaining part is a repair for the previous part’s damage.
5 · The shape it all adds up to
One lane per token, running straight up.
6 · Keep this card
The whole thing on one index card.
Why it exists
Ten messages into a chat, you paste a long block of text, discuss it for a while, and then type “shorten the second paragraph.” The model knows which paragraph. It reaches back past everything you’ve said since and picks the right one. Hold onto that — reaching backward across a long conversation is the running example for this post, and it’s precisely the thing the previous generation of sequence models was bad at.
Before transformers, the default way to handle a sequence — a sentence, a time series, an audio clip — was a RNN or its better-behaved cousin, the LSTM. You fed the model one token, it updated a hidden state, you fed it the next token, it updated again. Sentence as conveyor belt.
That design had two problems that compounded each other.
The first was parallelism. Step t needs the hidden state from t−1, which needs t−2, and so on. You couldn’t really use a GPU the way GPUs want to be used — thousands of cores at once — because the work was a chain. Buying more hardware didn’t help much, because the bottleneck was the dependency, not the throughput.
The second was long-range dependencies. By the time information from the start of a passage reached the end, it had been squeezed through hundreds of tiny update steps and mostly washed out. LSTMs helped but didn’t solve it. Connecting your “second paragraph” back to a block of text ten messages ago meant fighting the architecture — the information had to survive every intervening step of the conveyor belt without being overwritten.
The 2017 paper Attention Is All You Need (Vaswani et al.) made a sharper move than people expected: throw recurrence out entirely. Instead of carrying state forward step by step, let every token look directly at every other token, in parallel, in a single operation called self-attention. No conveyor belt: a distant token is one hop away instead of hundreds, so there’s nothing for it to get squeezed through. And the whole operation is embarrassingly parallel on a GPU.
That is the whole reason the transformer exists. The rest is engineering around that one swap.
Why it matters now
Once people had an architecture that scaled cleanly with compute, scale itself became the lever — and the transformer turned out to be ridiculously general. The same block design sits underneath the current LLM families (GPT, Claude, Gemini, Llama), the ViT line of vision models, and OpenAI’s Whisper for speech recognition. Attention-based blocks are also central to protein structure prediction — AlphaFold is built around them, though the full system is a good deal more than a stack of transformer blocks. If you’re trying to understand a current foundation model above the API layer, “transformer” is very likely the shape you’re looking at.
The short answer
transformer = stack of (self-attention + feed-forward) blocks, with positional info baked into the input
Picture to keep: one lane per token running straight up through the stack, and at every floor the lanes are allowed to reach sideways and copy from each other before each one goes off and thinks privately. Nothing is passed along the sequence; everything is written into a lane and read across. Where the picture breaks: the lanes aren’t sealed pipes carrying one token’s meaning, since after the first floor each lane holds a blend of whatever it read from the others.
A transformer turns a sequence of tokens into a sequence of contextual vectors by repeatedly mixing information across positions (self-attention) and then transforming each position individually (feed-forward). Stack that block enough times, train at scale, and you get the substrate behind today’s foundation models.
How it works
Take the one swap — no recurrence, every token looks at every other token — and follow what it forces. Each piece of the architecture is the repair for the previous piece’s damage.
Start: let every token look at every other token. Text is split into tokens (see tokenization) and each token ID looks up a vector in a learned table — its embedding, so a message becomes a matrix. Self-attention then lets each position build a weighted sum of every other position, weights learned and content-dependent. “Second paragraph” can read directly from the paragraph, in one step, however far back it sits. No conveyor belt to survive.
But now the model is order-blind. Shuffle the input vectors and self-attention gives you the same set of outputs, just shuffled to match — no position has any way to know it came first. “The dog bit the man” and “the man bit the dog” are the same bag of tokens, which is fatal for language. Fix: inject position into the input. The original paper used fixed sinusoidal positional encodings; many current decoder-only LLMs use rotary variants instead. See why positional encodings exist and why RoPE replaced sinusoidal.
But gathering isn’t thinking. What attention hands back is a weighted average of (linearly transformed) things already in the sequence — it’s very good at moving information sideways and poor at turning it into something new. Fix: after each attention step, run each position through a small MLP, independently. Gather, then think. That pair — attention plus feed-forward — is the transformer block, and you stack it N times so information can make several hops.
But a stack that deep is hard to train. Gradients pushed back through dozens of blocks tend to vanish or explode, and each block’s output can drift to a scale the next one wasn’t expecting. Fix: wrap every sublayer in a residual connection so the gradient has a straight path from top to bottom, and a layer norm to keep the scales in range. These two are not decoration — they’re most of the reason “just stack more blocks” became a viable strategy at all.
The residual connection is also what makes the picture-to-keep literal. Each token’s lane is the residual stream: a vector that travels straight up the stack, with every block adding into it. Attention reads and writes back; the FFN reads and writes back. The stream is the working memory.
Finally, turn a lane back into a word. At the top, each token’s vector is projected into a distribution over the vocabulary. For an LLM, the distribution over the last position is the next-token prediction. Sample, append, run the whole thing again — that’s generation.
I’m being deliberately loose about the math (queries, keys, values, softmax, head counts, hidden sizes). The shape is re-derivable; the hyperparameters change with every new model and aren’t worth memorizing.
There is a bill for all this, and it’s worth knowing you’re paying it: letting every token see every other token means N² pairs, so cost grows quadratically with sequence length — which is a post of its own, and most of why long context is expensive.
You started with transformer = stack of (attention + feed-forward) blocks + positional info. What did this post add? — + residuals and normalization, the least glamorous part of the list and most of the
reason the word stack is allowed to mean sixty layers instead of six.
Meanwhile the hook resolves in the first line of the mechanism: your
“second paragraph” is reachable because attention connects any two
positions directly, no matter how far apart. The depth on top of that
buys refinement, not reach — that’s my read of what the extra layers are
doing, not something the architecture states.
Famous related terms
- Attention —
attention = each token weights every other token by learned relevance. The mechanism the architecture is named after. - Self-attention vs cross-attention —
self-attention = a sequence attending to itself;cross-attention = one sequence attending to a different sequence(e.g. a decoder reading from an encoder’s output). - Positional encoding / RoPE —
positional encoding = "where am I in the sequence?" injected into the input. See also why RoPE replaced sinusoidal. - Decoder-only / encoder-only / encoder-decoder —
decoder-only = causal self-attention only;encoder-only = bidirectional self-attention only;encoder-decoder = encoder + decoder with cross-attention. Three flavors of the same block. Encoder-only (BERT-style) reads a whole sequence at once. Decoder-only (GPT-style) is causal — each position only attends to earlier ones, which is what makes next-token generation work. Encoder-decoder (the original 2017 design) does both. - Residual connection —
residual = output + input. Lets gradients flow straight from the top of the stack to the bottom, which is most of why very deep transformers train at all. - LLM —
LLM = transformer + "predict the next token" objective at scale. The most visible product of this architecture.
Going deeper
- Attention Is All You Need (Vaswani et al., 2017) — the primary source for “what exactly is a transformer block?”; section 3 has the precise definition of multi-head attention, and it’s short enough to read in one sitting.
- Jay Alammar, The Illustrated Transformer — the explainer for “what are the shapes actually doing?”, if the matrix dimensions in the paper won’t sit still in your head.
- Rabbit hole: Andrej Karpathy’s Neural Networks: Zero to Hero, and the Let’s build GPT video in it — answers “could I write one myself?” by building a tiny decoder-only transformer end to end.
A note on what I’m sure of: the high-level shape — embeddings, positional info, stacked attention + FFN blocks with residuals and norm, final projection — is the same across essentially every modern transformer. The details (head counts, hidden sizes, exact norm placement, positional scheme, FFN activation) vary model to model and aren’t pinned down here.