Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why RLHF exists

A pretrained language model knows everything and answers nothing. RLHF exists because the gap between 'predict the next token' and 'do what the user asked' is wider than prompt engineering can paper over.

AI & ML intermediate Apr 29, 2026 · updated Aug 25, 2026 · 14 min read

On this page

The picture version

Six pictures for a reader who has only ever met a chatbot from the outside. The prose below fills in the seams the pictures skip.

1 · The problem

Ask a freshly trained model a question and it doesn’t answer. It continues.

what is the capital of France? what is the capital of Germany? what is the capital of Spain? It isn’t broken. On the internet, that line is usually followed by more of the same. “likely continuation” and “helpful answer” are different targets
A model trained only to predict what comes next does exactly that — and a question is often followed by more questions. Answering is a different job from continuing, and nothing in the original training asked for it.

2 · Why you can’t just say what you want

Write down a score for “helpful.” Go on.

what you’d need helpful(answer) = ? nobody can write this down — not honestly, not in advance what people can do easily shown two answers: “that one’s better” reliably, quickly, thousands of times The whole method is built on that gap between the two boxes.
“Be helpful” is the kind of thing people recognise on sight and cannot specify in advance. The trick is to stop trying to write the rule and start collecting the judgements instead.

3 · First move: just show it

Thousands of worked examples. This does most of the visible work — and then stops.

questions, with good answers written out now it answers “Paris.” but you can’t write an example for every situation there is and you can’t demonstrate what not to do by doing it Showing good answers teaches the shape of an answer. It can’t teach preferences between two answers that both look fine.
Fine-tuning on human-written answers is cheap and does most of what you notice — the model stops continuing and starts answering. What it can’t do is cover the whole space, or express a negative, which is where the next stage comes in.

4 · The load-bearing trick

Hire a taste-tester: a second model, fitted to thousands of “this one’s better.”

two answers, same question … thousands of pairs ← better ← better fit the taste-tester reads an answer, returns one number a score you can actually optimise You haven’t defined “helpful.” You’ve curve-fitted people’s judgements of it. which is the difference that everything after this inherits — including the mistakes
Show people pairs, record which they prefer, and train a second model to predict those choices as a single number. An unwritable concept becomes a function you can push against — one fitted to a sample of people, with blind spots nobody listed.

5 · Practising against the palate — on a leash

Let the model chase that score, and it will find the holes in it.

what the taste-tester rewards, across all possible answers genuinely good scores even higher and is gibberish left alone, it walks this way so you tether it: climb, but don’t wander far from where you started
The scorer is a stand-in, not the real thing, and a model optimising hard against a stand-in will find its blind spots — answers that score beautifully and read as nonsense. A tether back to the starting model is what keeps the climb honest.

6 · Keep this card

The whole thing on one index card.

the recipe = show it thousands of good answers + fit a scorer to which of two people prefer + practise against the scorer, on a leash the middle line is the one the three-part name hides Nobody can write “helpful” down. People can pick the better of two. so the model’s manners have a shape — and it is the shape of whatever that fit rewards
Picture to keep: a taste-tester standing between the cook and the diner. Nobody can write the recipe for “good,” so you hire someone who has eaten thousands of pairs of dishes and can say which plate is better, and let the cook practise against that palate. Where it breaks, and it matters: the taste-tester is itself a model fitted to a sample of people, and the cook may practise against it indefinitely — which is where agreeing too readily, and sounding samey, come from.

Why it exists

Type a few words into your phone’s keyboard and it offers you the next one. It isn’t answering you — it’s continuing you. That is, at bottom, the same objective a LLM is trained on, just at an incomprehensibly larger scale. Which raises a question worth sitting with: if the training objective is “guess what comes next,” why does a chatbot answer your question instead of just continuing it?

It doesn’t, at first. Take a freshly pretrained LLM — one that has only seen the next-token objective on a giant pile of internet text — and ask it “what is the capital of France?” and you are not guaranteed to get “Paris.” You might get a plausible continuation like “what is the capital of Germany? what is the capital of Spain?” — because in the training data, that question often shows up in a list of similar questions.

The model isn’t broken. It’s doing exactly what it was trained to do: predict likely next tokens given the prompt. The problem is that “likely continuation of this string on the internet” and “answer this question helpfully” are different distributions, and the second one is the one users want.

You can paper over this with prompt engineering — few-shot examples, instruction phrasings, formatting tricks — and that worked for a while. But the gap is too wide and too varied to close with prompts alone. The model needs to be trained to prefer the helpful continuation over the merely likely one. The natural move is to write down a hand-coded reward function — “give +1 for helpful, −1 for unhelpful” — gradient-descend on it, and call it a day. The catch: nobody has a clean hand-written function for “helpful.” It’s the kind of thing humans recognize when they see it but cannot specify in advance.

RLHF exists because that’s exactly what RL was built for: optimizing against rewards you can score but can’t hand-write. The “reward” here is still a function — but a learned one, a second model trained on humans comparing pairs of outputs. The LLM is then nudged, via reinforcement learning, to produce outputs that the reward model rates highly. That whole stack is the foundation of what “alignment” means in practice for the current generation of chatbots, even as the techniques on top of it have evolved.

Why it matters now

The chat assistants people actually use are post-RLHF (or RLHF-adjacent) models. The base pretrained checkpoint behaves very differently from the product: InstructGPT’s own evaluations found human raters preferred the 1.3B post-trained model’s outputs to the 175B base model’s, which is the cleanest published statement of how much of the “assistant” is post-training rather than scale.

What RLHF (and its successors — DPO, RLAIF, constitutional methods, and the reasoning-model RL recipes) buy you in production:

For an engineer building on top of these models, the practical consequence is that much of what you observe at the API — refusal patterns, formatting habits, verbosity, the way it handles ambiguous instructions — is post-training-shaped, not pretraining-shaped. (Not all of that post-training is literally RLHF in 2026; SFT, DPO-family methods, AI-feedback variants, and reasoning-model RL all live in the same stage.) When the model surprises you, the surprise often lives in post-training, not the base.

The short answer

RLHF = supervised fine-tuning + reward model trained on human preference comparisons + RL loop that optimizes the LLM against the reward model

Picture to keep: a taste-tester standing between the cook and the diner. Nobody can write down the recipe for “good,” so you hire someone who has eaten thousands of pairs of dishes and can reliably say which of two plates is better — and then you let the cook practice against that person’s palate, not against the diner’s. Where the analogy breaks, and it’s the important part: the taste-tester here isn’t a person, it’s a model fit to a sample of people’s judgments. It has blind spots nobody listed, and the cook is allowed to practice against it indefinitely.

You can’t write a loss function for “be helpful,” so you train a second model to predict which of two responses a human would prefer, and then use reinforcement learning to push the LLM toward outputs the second model rates highly. The reward model is the proxy; the RL loop is how you cash that proxy into weight updates.

How it works

The canonical recipe — the one OpenAI used for InstructGPT (Ouyang et al., 2022, arXiv:2203.02155) and that became the default template for the field — has three stages. They’re easiest to remember as three attempts at the same problem, each added because the previous one ran out of road.

Stage 1 — Supervised fine-tuning (SFT)

The obvious first move: if the model continues questions instead of answering them, show it thousands of examples of questions being answered. Start with the pretrained base model, collect a modest dataset of prompts paired with high-quality human-written responses, and fine-tune with the ordinary next-token objective. This is the cheapest, simplest step and it does most of the visible work: after SFT alone, “what is the capital of France?” gets you Paris, and the model already feels much more like an assistant. Most of “instruction following” is already happening here.

SFT alone has limits, though. You can only show the model so many human-written examples, and “good response” is high-dimensional — the SFT data captures one cross-section of it, not the whole space. You also can’t easily teach negative preferences (don’t be confidently wrong, don’t pad) by showing more positive examples.

Stage 2 — Train a reward model on preference pairs

So write more examples? That’s where the road runs out: you’d need a demonstration for every situation, and you still couldn’t demonstrate what not to do. Take the SFT model. For a batch of prompts — say twenty different phrasings of geography questions like our capital-of-France one — sample several responses from it. Show pairs to human raters and ask: which one is better? Train a reward model (often initialized from the LLM itself, with the language-modeling head replaced by a scalar score head) to predict the human’s choice. InstructGPT uses the Bradley-Terry preference loss — for a pair (chosen, rejected), maximize the log-probability that the reward of “chosen” exceeds the reward of “rejected.”

This is the load-bearing trick of the whole pipeline. You’ve turned an unspecified concept (“helpful”) into a learned function — one that takes a prompt and a candidate response and returns a number you can differentiate against. You haven’t defined helpfulness; you’ve curve-fit human judgments of it.

Stage 3 — RL against the reward model

A scorer alone changes nothing — it can rank the model’s answers but not improve them. Now use reinforcement learning to update the LLM’s weights so its sampled outputs score higher under the reward model. The InstructGPT paper used PPO, and because that paper was the template, PPO is what most early RLHF work reached for (the landscape has since diversified — DPO, GRPO and others). To keep the model from drifting into bizarre high-reward regions, the loss includes a KL penalty against the SFT model: “go up the reward gradient, but don’t move too far from where you started.”

Without that KL leash, you get reward hacking: the LLM finds outputs that score high on the reward model but a human would call gibberish. The reward model is a proxy, and the LLM is good at finding holes in proxies.

Why “RL” specifically?

This is the part that confused everyone (including me) at first. If you have a differentiable reward model, why not just backprop through it? The honest answer: there’s no clean gradient path. The LLM produces tokens by sampling, and the sampling step (argmax / categorical draw over the vocab) is non-differentiable. You can finesse this with relaxations or sequence-level objectives, but policy-gradient RL is the well-trodden workaround: PPO and its cousins push the LLM’s distribution in directions that, on average, raise the reward, without needing a gradient through the sampling step.

The 2017 paper that put preference-based RL on the map for deep learning is Christiano et al., Deep Reinforcement Learning from Human Preferences, which trained agents on Atari and simulated locomotion using human comparisons of trajectory pairs. The InstructGPT line is the application of that idea to language models.

What goes wrong, and what came after

A few seams worth knowing:

You started with RLHF = SFT + reward model + RL loop. What did this post add that the three-part name hides? — + the reason the middle piece has to be *learned*. Nobody can write “helpful” as a function, but people can reliably pick the better of two answers, so you curve-fit their picks and optimize against the fit. That’s what turns your phone’s autocomplete into something that answers “what is the capital of France?” instead of continuing it — and it’s also why the model’s personality has a shape: everything downstream inherits the quirks of whatever that curve-fit rewards.

Check yourself

Before you go — you have a great reward model and a fast GPU cluster, so you run stage 3 for far longer than usual, watching the average reward climb the whole time. Human evaluators then rate the result worse than the model you started with. What happened, and which single term in the objective was supposed to prevent it?

Answer

Rising reward means the policy is finding outputs the reward model likes, which is not the same as outputs humans like — the reward model is a curve fit to a slice of preference data, and long optimization pushes into regions that slice never covered. That’s reward hacking. The term meant to bound it is the KL penalty against the SFT model: it prices movement away from the starting distribution, so the policy can’t wander arbitrarily far to chase reward. Note the trade it forces — turn KL up and you learn less; turn it down and you drift more. There’s no setting that makes the mismatch disappear.

And one more: if the reward model is differentiable, why not just backpropagate the reward into the LLM and skip reinforcement learning entirely?

Answer

Because of what sits between the LLM and the reward model: sampling. The LLM outputs a distribution, and turning that into the actual token sequence the reward model scores is a discrete draw, which has no useful gradient. Policy-gradient methods like PPO get around this by estimating which direction to nudge the distribution from sampled outcomes rather than differentiating through them. And this also explains why DPO caused such a stir: it rewrites the objective so that during fine-tuning you train directly on the stored preference pairs with a classification-style loss — no sampling inside the training loop to route around.

Going deeper

What’s well documented: the three-stage recipe, the reason RL is used (sampling is non-differentiable), and the failure modes (reward hacking, sycophancy, mode collapse). What isn’t: the exact post-training recipes frontier labs run in 2026. The public papers describe the shape of the pipeline, but the data mixtures, the reward-model architectures, and how PPO-vs-DPO-vs-something-else has shaken out at scale are mostly proprietary.