Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why does in-context learning work?

You paste three examples into a prompt and the model suddenly does the task. Nothing got trained. So what just happened?

AI & ML intermediate Apr 29, 2026 · updated Aug 25, 2026 · 13 min read

On this page

The picture version

Six pictures for a reader who has never thought about what a prompt does. The prose below fills in the seams the pictures skip.

1 · The thing that happens

Three examples, and it finishes the fourth.

what the spreadsheet says what you want APPLE IPHONE 13 PRO- 128gb iPhone 13 Pro 128GB samsung galaxy s22 (blk) Galaxy S22 Black GOOGLE Pixel_7a -128 Pixel 7a 128GB SONY wh1000xm4/blk WH-1000XM4 Black you never wrote down the rule — you just showed three of them
You paste three before-and-after pairs into a chat box, add a fourth messy name with nothing after it, and the right answer comes back. No rule was written and no file was uploaded.

2 · The wrong picture

Nothing was taught. Nothing about the model changed.

what it feels like your 3 examples teach the model’s numbers nudged a little no learning step runs. no numbers move. nothing is written back anywhere. what happens frozen your 3 examples go here instead the prompt — just text the model reads
The weights — the numbers fixed during training — are identical before and after. Run the same prompt against a read-only copy on another machine and you get the same behaviour. Whatever happened, happened inside one pass of reading your text.

3 · The mechanism

Your examples are just earlier words in the same sentence.

example 1 example 2 example 3 the new messy name ? messy → clean messy → clean messy → clean the position that has to produce the answer can read every earlier position one pass · the same frozen numbers · nothing kept afterwards “few-shot learning” is, mechanically, just a longer piece of text
There is no separate learning phase to point at. The same arithmetic that ran on the first word runs on the five-thousandth — the examples influence the answer the way any earlier words in a sentence do.

4 · The key idea

The examples don’t add the ability. They pick it out.

your three examples, lying across the strings still ringing everything the model can already do, all at once dashed strings: damped silent  ·  left ringing: clean up product names
A piano that plays every piece it has ever heard, badly, all at once. Your examples are a hand damping every string except the ones belonging to the task. The ability was already in the instrument — though which structures inside the model actually get selected is still an open research question.

5 · The three tells

Stress it, and it behaves nothing like training.

reorder the same three examples near-random → near-best a training set doesn’t care what order it’s written in. this does. make the answers in the examples wrong barely a dent barely moves. so it isn’t fitting your pairs — it’s reading the format and the shape. close the tab and come back gone nothing persisted, because nothing was ever written. a new chat starts blank. the first two panels are measured results — the studies are cited in the prose below
All three are bizarre if your examples taught the model something, and unsurprising if they selected something it already had. The seams are the evidence for the picture, not exceptions to it.

6 · Keep this card

The whole idea on one index card.

in-context learning = your examples in the prompt + numbers that never change + behaviour that mimics being trained — “mimics” is the load-bearing word
Picture to keep: a piano that plays every piece it has ever heard, badly, all at once — and your three examples are the hand damping every string but one set.

Why it exists

You have a spreadsheet of messy product names to clean up. You paste three before-and-after examples into a chat box, then a fourth messy name with nothing after it, and the model finishes it correctly. You never wrote down the rule. You’ve probably done some version of this and thought nothing of it.

That spreadsheet is the running example for this post, and here is its canonical published form — the shape LLMs were shown to do in the paper that introduced GPT-3 (Brown et al., 2020):

Translate English to French.
sea otter => loutre de mer
peppermint => menthe poivrée
plush girafe =>

Nothing was trained here. No weights were updated, no fine-tuning happened, and the model has no memory of the previous line by the time you close the tab. You typed some text into a box and got a French phrase back. You probably assume something got taught to the model in that moment. It’s the opposite: nothing about the model changed at all — and holding onto the “it learned” picture is what makes every failure mode in this post look like a bug instead of a clue.

That same paper is what put the name in-context learning into general circulation — and the field spent the next several years trying to figure out what it actually is. The honest answer is that we still don’t fully know. There are good partial stories. None of them is settled.

I am writing this post because the curious-engineer question — the weights didn’t change, so what kind of “learning” is this? — is more interesting than any how-to about prompting.

Why it matters now

Take in-context learning away and a lot of what modern LLM products do would have to be rebuilt as training runs instead of prompts.

The whole stack assumes a model that adapts from context. If you don’t know why that works, you can’t predict when it will fail.

The short answer

in-context learning = examples in the prompt + a model whose weights never change + behaviour that mimics having been trained on them

Picture to keep: a piano that plays every piece it has ever heard, badly, all at once. Your three examples aren’t teaching it a new piece — they’re a hand on the keyboard damping every string except the ones that belong to clean up product names. Except that no one has found the strings: the damping is real and measurable, but which internal structures your examples actually select is exactly the open question in the second half of this post.

Nothing literally learns at inference. The weights are frozen. What changes is the model’s conditional distribution over the next token once it has been forced to attend to your examples. Because the training objective happened to make those conditional distributions behave a lot like “do the task the examples are doing,” it looks like the model picked up a skill. It didn’t — the capability, in whatever partial and fragile form, was already latent in the weights, and your prompt selected it.

How it works

The obvious explanation is that the model trains on your examples on the fly — a few gradient steps, then back to inference. Why that’s wrong: inference is a pure function of the weights and the input. There is no optimizer in the loop, no gradient computed, nothing written back. Run the same prompt on a read-only copy of the weights on a different machine and you get the same behaviour. So whatever is happening is happening in the activations of a single forward pass, and two questions fall out that have very different answers:

  1. At inference time, what is the model mechanically doing?
  2. Why did training on next-token prediction give it that ability in the first place?

What’s mechanically happening

Inside the transformer, your prompt becomes a sequence of token embeddings. Each layer’s attention mechanism lets later positions read from earlier ones. By the time the model is computing the distribution for the next token, every previous token in the prompt — including your three “sea otter ⇒ loutre de mer” examples — has had a chance to influence the internal state.

So “few-shot learning” is, at the level of the math, just a longer context window. There is no parameter update, no gradient, no separate “learning phase.” The same matrix multiplications that ran on token 1 run on token 5,000. The weights never know they’re doing translation. The activations do.

That reframing is useful: in-context learning is whatever the attention pattern does when it conditions on patterned context. It’s a property of inference, not a separate algorithm.

Why next-token training gave us this

This is the part that is not fully understood, and I want to be clear about that. Several partial accounts have real evidence behind them; none of them is “the answer.”

The “implicit Bayesian inference” story. Xie et al. (2022) and others argued that a model pretrained on text with hidden structure learns to infer which latent task a passage is doing — they show this in a synthetic setting, and the extension to real corpora is the intuition rather than the result. The everyday version of that intuition: natural text is full of stretches that look like a task followed by examples of it — recipes, FAQs, code with docstrings, exam answers, translation pairs. The model learns to infer “which task is this passage doing?” as a side-effect of next-token prediction, because guessing the task helps predict the next token. At inference, your few-shot prompt looks like one of those stretches; the model infers the task from your examples and generates accordingly. In this view, in-context learning isn’t a new capability — it’s the model doing the same task-inference it always did, with your prompt as the input.

The “induction head” story. Olsson et al. (2022, Anthropic) found small circuits inside transformers — pairs of attention heads they called induction heads — that implement a very specific behavior: if the pattern [A][B] appeared earlier in the context, and you now see [A] again, attend to and copy [B]. They showed these circuits form abruptly during training, and the moment they form coincides with the model getting much better at in-context learning on synthetic tasks. The claim is not that induction heads explain all in-context learning, but that they’re a concrete, mechanistic example of how next-token training can build a circuit that looks, from outside, like “learning from examples.”

The “gradient descent in the forward pass” story. Garg et al. (2022) showed that transformers can learn simple function classes — linear regression and friends — purely from in-context (input, output) pairs. von Oswald et al. (2022) and Akyürek et al. (2022) then argued the stronger claim: on those tasks the model’s predictions closely match what one or a few steps of gradient descent on those pairs would produce. The provocative reading: the forward pass is implementing a tiny optimizer over the in-context examples. How far this generalizes from toy regression to real natural language is contested. It’s a beautiful mathematical result and an uncertain empirical claim.

These stories are not mutually exclusive. They’re probably all partially right, on different tasks, at different scales. The honest summary is: we have several mechanistic hypotheses with supporting evidence, and we do not yet have a unified theory of why scaling a next-token predictor produces this behavior. Anyone who tells you otherwise with confidence is overselling.

Where the seams show

If you stress in-context learning, it cracks in instructive ways:

These are not bugs to fix. They’re tells that whatever in-context learning is, it’s not the same kind of process as training. It’s a sibling, not a copy — and notice that every one of them is bizarre under “the model learned French from your three pairs” and unsurprising under “your three pairs selected a behaviour the model already had.”

You started with in-context learning = examples in the prompt + a model whose weights never change + behaviour that mimics having been trained. What did the seams add? — + "mimics" is load-bearing. The examples are doing more selecting than teaching; that one substitution predicts the order sensitivity, the random-label result, the flattening returns, and the fact that closing the tab erases it.

Check yourself

Before you go — your product-name cleanup is getting 71% right with four in-context examples. You spot that one of your four examples has the wrong output, fix it, and accuracy barely moves. Then you swap the order of two examples and it jumps to 84%. What does that pattern tell you about what your examples are doing?

Answer

That they’re mostly specifying format and label space, not teaching the input→label mapping — which is what Min et al. (2022) found on the classification-style tasks they tested, where replacing labels with random ones did little damage. If the model were fitting your demonstrations, a corrected label would matter and ordering wouldn’t. That it’s the other way round says the examples are selecting a behaviour the model already has. Practical upshot for a task shaped like this one: spend your effort on format, label vocabulary, and example ordering before you spend it on label-checking — and measure ordering, because Lu et al. (2022) showed it can swing results across nearly the whole achievable range.

And: someone proposes fixing the remaining 29% by cramming 500 cleaned product names into a long context window instead of fine-tuning on them. Under the model in this post, what would you predict, and what’s the one thing you’d measure to check?

Answer

Predict sharply diminishing returns — the first handful of examples do most of the work, and there’s no general law saying the next 450 keep paying, which is unlike what more training data buys a real learning algorithm. The thing to measure is whether accuracy on a held-out set keeps improving from 50 → 200 → 500 examples, and separately whether it’s sensitive to which 500 and in what order. If shuffling the 500 changes the score, you have selection, not learning, and more examples aren’t the lever. (Fine-tuning does update weights, so it has a route to keep improving past the point where prompt-stuffing flattens out — though whether it actually will on your data is its own question.)

Going deeper

Where the line sits: the phenomenon — models behaving as if they learned from in-prompt examples, with no weight updates — is rock-solid and reproducible. The explanation is genuinely open research, which is why the three stories above are labelled as stories and the evidence for each is scoped to the setting it was measured in.