Why does in-context learning work?
You paste three examples into a prompt and the model suddenly does the task. Nothing got trained. So what just happened?
On this page
The picture version
Six pictures for a reader who has never thought about what a prompt does. The prose below fills in the seams the pictures skip.
1 · The thing that happens
Three examples, and it finishes the fourth.
2 · The wrong picture
Nothing was taught. Nothing about the model changed.
3 · The mechanism
Your examples are just earlier words in the same sentence.
4 · The key idea
The examples don’t add the ability. They pick it out.
5 · The three tells
Stress it, and it behaves nothing like training.
6 · Keep this card
The whole idea on one index card.
Why it exists
You have a spreadsheet of messy product names to clean up. You paste three before-and-after examples into a chat box, then a fourth messy name with nothing after it, and the model finishes it correctly. You never wrote down the rule. You’ve probably done some version of this and thought nothing of it.
That spreadsheet is the running example for this post, and here is its canonical published form — the shape LLMs were shown to do in the paper that introduced GPT-3 (Brown et al., 2020):
Translate English to French.
sea otter => loutre de mer
peppermint => menthe poivrée
plush girafe =>
Nothing was trained here. No weights were updated, no fine-tuning happened, and the model has no memory of the previous line by the time you close the tab. You typed some text into a box and got a French phrase back. You probably assume something got taught to the model in that moment. It’s the opposite: nothing about the model changed at all — and holding onto the “it learned” picture is what makes every failure mode in this post look like a bug instead of a clue.
That same paper is what put the name in-context learning into general circulation — and the field spent the next several years trying to figure out what it actually is. The honest answer is that we still don’t fully know. There are good partial stories. None of them is settled.
I am writing this post because the curious-engineer question — the weights didn’t change, so what kind of “learning” is this? — is more interesting than any how-to about prompting.
Why it matters now
Take in-context learning away and a lot of what modern LLM products do would have to be rebuilt as training runs instead of prompts.
- Prompt engineering is in-context learning by another name. Every “you are a helpful assistant who…” preamble is leaning on the model’s ability to specialize on the fly.
- Few-shot prompting — your spreadsheet cleanup — is usually the first thing people try before reaching for fine-tuning, and often it works well enough that fine-tuning never gets built.
- RAG leans on the same conditioning: drop retrieved passages into the prompt and the model answers from them. It’s not few-shot task inference exactly — the documents are content, not demonstrations — but it’s the same underlying property that the prompt can steer behaviour without touching weights.
- Tool use, agents, structured output lean on it too: show the model the shape and it tends to produce that shape. (Instruction tuning and, for structured output, constrained decoding do a lot of the work here as well — this isn’t in-context learning alone.)
The whole stack assumes a model that adapts from context. If you don’t know why that works, you can’t predict when it will fail.
The short answer
in-context learning = examples in the prompt + a model whose weights never change + behaviour that mimics having been trained on them
Picture to keep: a piano that plays every piece it has ever heard, badly, all at once. Your three examples aren’t teaching it a new piece — they’re a hand on the keyboard damping every string except the ones that belong to clean up product names. Except that no one has found the strings: the damping is real and measurable, but which internal structures your examples actually select is exactly the open question in the second half of this post.
Nothing literally learns at inference. The weights are frozen. What changes is the model’s conditional distribution over the next token once it has been forced to attend to your examples. Because the training objective happened to make those conditional distributions behave a lot like “do the task the examples are doing,” it looks like the model picked up a skill. It didn’t — the capability, in whatever partial and fragile form, was already latent in the weights, and your prompt selected it.
How it works
The obvious explanation is that the model trains on your examples on the fly — a few gradient steps, then back to inference. Why that’s wrong: inference is a pure function of the weights and the input. There is no optimizer in the loop, no gradient computed, nothing written back. Run the same prompt on a read-only copy of the weights on a different machine and you get the same behaviour. So whatever is happening is happening in the activations of a single forward pass, and two questions fall out that have very different answers:
- At inference time, what is the model mechanically doing?
- Why did training on next-token prediction give it that ability in the first place?
What’s mechanically happening
Inside the transformer, your prompt becomes a sequence of token embeddings. Each layer’s attention mechanism lets later positions read from earlier ones. By the time the model is computing the distribution for the next token, every previous token in the prompt — including your three “sea otter ⇒ loutre de mer” examples — has had a chance to influence the internal state.
So “few-shot learning” is, at the level of the math, just a longer context window. There is no parameter update, no gradient, no separate “learning phase.” The same matrix multiplications that ran on token 1 run on token 5,000. The weights never know they’re doing translation. The activations do.
That reframing is useful: in-context learning is whatever the attention pattern does when it conditions on patterned context. It’s a property of inference, not a separate algorithm.
Why next-token training gave us this
This is the part that is not fully understood, and I want to be clear about that. Several partial accounts have real evidence behind them; none of them is “the answer.”
The “implicit Bayesian inference” story. Xie et al. (2022) and others argued that a model pretrained on text with hidden structure learns to infer which latent task a passage is doing — they show this in a synthetic setting, and the extension to real corpora is the intuition rather than the result. The everyday version of that intuition: natural text is full of stretches that look like a task followed by examples of it — recipes, FAQs, code with docstrings, exam answers, translation pairs. The model learns to infer “which task is this passage doing?” as a side-effect of next-token prediction, because guessing the task helps predict the next token. At inference, your few-shot prompt looks like one of those stretches; the model infers the task from your examples and generates accordingly. In this view, in-context learning isn’t a new capability — it’s the model doing the same task-inference it always did, with your prompt as the input.
The “induction head” story.
Olsson et al. (2022, Anthropic) found small circuits inside
transformers — pairs of attention heads they called induction heads —
that implement a very specific behavior: if the pattern [A][B]
appeared earlier in the context, and you now see [A] again, attend
to and copy [B]. They showed these circuits form abruptly during
training, and the moment they form coincides with the model getting
much better at in-context learning on synthetic tasks. The claim is
not that induction heads explain all in-context learning, but that
they’re a concrete, mechanistic example of how next-token training can
build a circuit that looks, from outside, like “learning from
examples.”
The “gradient descent in the forward pass” story. Garg et al. (2022) showed that transformers can learn simple function classes — linear regression and friends — purely from in-context (input, output) pairs. von Oswald et al. (2022) and Akyürek et al. (2022) then argued the stronger claim: on those tasks the model’s predictions closely match what one or a few steps of gradient descent on those pairs would produce. The provocative reading: the forward pass is implementing a tiny optimizer over the in-context examples. How far this generalizes from toy regression to real natural language is contested. It’s a beautiful mathematical result and an uncertain empirical claim.
These stories are not mutually exclusive. They’re probably all partially right, on different tasks, at different scales. The honest summary is: we have several mechanistic hypotheses with supporting evidence, and we do not yet have a unified theory of why scaling a next-token predictor produces this behavior. Anyone who tells you otherwise with confidence is overselling.
Where the seams show
If you stress in-context learning, it cracks in instructive ways:
- Order matters. Lu et al. (2022) showed that just permuting the order of few-shot examples can move the same model on the same task anywhere from near-random to near-state-of-the-art. A standard supervised objective doesn’t care what order a dataset is written down in — the loss is a sum over examples. This does.
- The labels barely have to be right. Min et al. (2022) found that on many classification tasks, replacing the example labels with random labels barely hurt few-shot performance — what mattered was the format and the label space, not the input-label mapping. This is hard to square with “the model is learning the task from the examples.” It’s much easier to square with “the examples are telling the model which distribution of behavior to switch into.”
- The returns flatten, and where they flatten is task-specific. The first few examples help sharply; after that the curve depends on the task, the model, and how long the context is, and there is no general law saying more examples keep paying. That is itself the tell: with a real learning algorithm, more training data reliably buys more, and here it doesn’t.
- It’s not durable. End the conversation, start a new one, and the “skill” is gone. Whatever happened was scoped to the activations of one forward pass.
These are not bugs to fix. They’re tells that whatever in-context learning is, it’s not the same kind of process as training. It’s a sibling, not a copy — and notice that every one of them is bizarre under “the model learned French from your three pairs” and unsurprising under “your three pairs selected a behaviour the model already had.”
You started with in-context learning = examples in the prompt + a model whose weights never change + behaviour that mimics having been trained.
What did the seams add? — + "mimics" is load-bearing. The examples are
doing more selecting than teaching; that one substitution predicts the
order sensitivity, the random-label result, the flattening returns, and
the fact that closing the tab erases it.
Check yourself
Before you go — your product-name cleanup is getting 71% right with four in-context examples. You spot that one of your four examples has the wrong output, fix it, and accuracy barely moves. Then you swap the order of two examples and it jumps to 84%. What does that pattern tell you about what your examples are doing?
Answer
That they’re mostly specifying format and label space, not teaching the input→label mapping — which is what Min et al. (2022) found on the classification-style tasks they tested, where replacing labels with random ones did little damage. If the model were fitting your demonstrations, a corrected label would matter and ordering wouldn’t. That it’s the other way round says the examples are selecting a behaviour the model already has. Practical upshot for a task shaped like this one: spend your effort on format, label vocabulary, and example ordering before you spend it on label-checking — and measure ordering, because Lu et al. (2022) showed it can swing results across nearly the whole achievable range.
And: someone proposes fixing the remaining 29% by cramming 500 cleaned product names into a long context window instead of fine-tuning on them. Under the model in this post, what would you predict, and what’s the one thing you’d measure to check?
Answer
Predict sharply diminishing returns — the first handful of examples do most of the work, and there’s no general law saying the next 450 keep paying, which is unlike what more training data buys a real learning algorithm. The thing to measure is whether accuracy on a held-out set keeps improving from 50 → 200 → 500 examples, and separately whether it’s sensitive to which 500 and in what order. If shuffling the 500 changes the score, you have selection, not learning, and more examples aren’t the lever. (Fine-tuning does update weights, so it has a route to keep improving past the point where prompt-stuffing flattens out — though whether it actually will on your data is its own question.)
Famous related terms
- Few-shot prompting —
few-shot = task description + k worked examples + the new input. The most common shape of in-context learning in practice. - Zero-shot prompting —
zero-shot = task description + the new input, no examples. Leans on the task already being latent from pretraining — and, for any chat model you’d actually use, on instruction tuning having taught it to treat a description as a request. - Chain of thought —
CoT ≈ in-context learning + 'show your work' as the demonstrated format. Same selection move, with reasoning steps as the thing being selected. - Fine-tuning —
fine-tuning = more training + your data + actual weight updates. The thing in-context learning lets you skip most of the time — and the thing you fall back to when the plateau bites. - Induction head —
induction head ≈ attention heads that copy "what came after [A] last time" when they see [A] again. Olsson et al. describe it as a two-head circuit in the models they studied; treat that as a well-characterized instance, not a universal definition. Still the cleanest mechanistic example of an in-context-learning building block. - Prompt engineering —
prompt engineering = exploiting in-context learning by hand. The applied side of all of the above.
Going deeper
- Language Models are Few-Shot Learners (Brown et al., 2020) — the primary source: where the sea-otter prompt comes from, and how the effect scales with model size.
- In-context Learning and Induction Heads (Olsson et al., Anthropic, 2022) — the best answer available to “can you actually point at the circuit doing this?”, with the training-time phase change as evidence.
- Rethinking the Role of Demonstrations (Min et al., 2022) — the rabbit hole: pull this thread if you want to know how little of your prompt’s content the model is really using.
Where the line sits: the phenomenon — models behaving as if they learned from in-prompt examples, with no weight updates — is rock-solid and reproducible. The explanation is genuinely open research, which is why the three stories above are labelled as stories and the evidence for each is scoped to the setting it was measured in.