Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

RAG: why retrieval didn't die when context windows got huge

Long context windows were supposed to kill retrieval-augmented generation. They didn't. Here's why the bottleneck moved instead of disappearing.

AI & ML intro Apr 29, 2026 · updated Aug 25, 2026 · 10 min read

On this page

The picture version

Six pictures for a reader who has never heard the term. The prose below fills in the seams the pictures skip.

1 · The problem

Fluent, instant, and about somebody else’s company.

“how much parental leave do I get?” “You’re entitled to 12 weeks of paid parental leave, with the option to…” confident. immediate. wrong. your handbook — never opened other people’s documents it answered from The model has read a great deal. None of it was yours.
Everything the model knows came from its training text, which stopped at some date and never included your internal documents. Asked about them, it produces a plausible answer instead of no answer — in exactly the same tone as a correct one.

2 · First obvious fix

Train it on the handbook. Now it answers — and you still can’t check it.

the handbook train the model, now stirred in “You get 16 weeks.” — from where? no answer. HR edits the policy → retrain the document and the model drift apart no paragraph to click so you still can’t tell right from invented
Folding the facts into the model can work. What it doesn’t give you is a link back to the sentence the number came from — and that receipt was the thing you actually wanted, because the failure in scene 1 was indistinguishable from success.

3 · Second obvious fix

Paste in everything. Pay for it, every single message.

the bill 500,000 words in, for a 12-word question × every message a big window is a capacity, not a free lunch the middle sags start end middle how reliably the fact gets used, by where it sits in the prompt it keeps changing HR edits the handbook new tickets arrive the code gets a commit an index updates one document. a giant paste updates nothing. A bigger window moved the bottleneck. It didn’t remove it. the curve above is the shape of a measured effect, not a specific model’s numbers — the study is cited in the prose
Three independent costs, none of which a larger context window fixes: you pay per message, the model uses the middle of a long prompt less reliably than its ends, and a pasted corpus goes stale the moment anyone edits anything.

4 · The fix

Fetch four pages before answering, not the whole library.

done once, in advance — and re-done for one document when it changes all your documents cut into passages filed by what it’s about the index the question look up the few passages that fit 4 passages ≈ 4,000 words the model 4,000 words per question instead of 500,000 — and the answer can quote the paragraph it came from.
The library stays outside the model and only a few pages come in per question. That fixes all three costs at once: a small prompt, a short context the model handles well, and an index you can update one document at a time.

5 · The seam

The librarian never reads the pages it hands you.

what came back parental leave — eligibility parental leave — the 2019 policy leave: how to apply sabbatical policy nothing flags the stale one matching happens on similarity, not truth the model answers just as fluently from the wrong page the two defences a second, slower pass that actually reads question and passage together + a citation you can click Retrieval doesn’t remove the wrong-answer problem. It makes it checkable.
“Looks relevant” is doing enormous work here, and it fails quietly — wrong pages arrive with no signal that they’re wrong. Ranking is where most of the accuracy is won; the citation is what lets a human catch the rest.

6 · Keep this card

The whole idea on one index card.

RAG = fetch the few passages that fit + answer from them, not from memory + keep the receipt — the ranking step is where the accuracy lives
Picture to keep: a librarian who pulls four pages before you ask, rather than handing you the whole library — except this librarian never reads them, so the wrong four pages arrive silently.

Why it exists

You ask the company chatbot “how much parental leave do I get?” and it answers immediately, in fluent HR prose, with a number. The number is wrong. It has never read your employee handbook. It read a great many other companies’ documents during training, and what came out is a plausible policy rather than yours. That handbook is the running example for this post.

That’s the wall apps built on top of an LLM ran into once they moved past demos: the model didn’t know your data. It knew a slice of the open web up to its training cutoff, and that was it. Your codebase, your customer’s tickets, your internal wiki, last week’s Slack — invisible.

The obvious fix was “just put the handbook in the prompt.” But the amount of text a model can read at once — its context window — was a few thousand words in 2022, and the handbook plus every other internal doc was a few million. So people did the next thing: cut the documents into passages, turn each passage into a list of numbers that captures roughly what it’s about, store those, and when a question arrives fetch only the passages that look relevant. Stuff those into the prompt. That’s Retrieval-Augmented Generation, and it became the standard answer to “how do I do AI on my own data?”

Then context windows grew, fast: 100k tokens in mid-2023, and a million with Gemini 1.5 in early 2024. And a recurring take appeared: RAG is dead, just paste the whole corpus in.

It didn’t happen. Retrieval didn’t go away, and the interesting question is why.

Why it matters now

If you’re building anything that talks to an LLM about non-public information — a coding agent, a support bot, a search-over-docs feature, a notes app — you’ll almost certainly end up with some retrieval step, even if you arrive at it by accident and call it something else. Understanding why the bottleneck moved instead of disappearing is the difference between a system that works at 100 documents and one that works at 100 million.

The short answer

RAG = retriever + generator + prompt assembly

Picture to keep: a librarian who pulls four pages for you before you ask your question, rather than handing you the whole library — or expecting you to have memorized it. Except that this librarian never reads the pages: it matches on shape and similarity, so the wrong four pages arrive silently, and the model answers from them just as fluently.

A retriever picks a small number of relevant chunks from a big corpus, an LLM generates an answer conditioned on them, and a thin layer in between decides what actually goes into the prompt. The reason huge context windows didn’t kill it: putting more tokens in the prompt costs money and latency on every turn, and even when you can afford it, the model gets worse at using them.

How it works

Take the parental-leave question and try to fix it the obvious ways.

Naive attempt: fine-tune the model on your handbook. Teach it the facts directly. Why it breaks: you have to redo it whenever HR edits the policy, and — the killer — the answer comes back with no provenance. The model can’t show you the paragraph it got “16 weeks” from, so you’re back to not knowing whether it’s real. Further training can teach facts; what it doesn’t give you is an inspectable link back to the source text, which is the thing you actually needed here. (A fine-tuned model can certainly print something citation-shaped. Whether that citation points at a real paragraph is exactly the question you were trying to settle.)

Fix: put the handbook in the prompt. Now the answer is grounded in text the model can quote, and updating the policy means updating a document. This is the right shape. Why it breaks at scale: three separate ways.

Fix: retrieve first, then prompt. Chunk the corpus, embed each chunk, and at query time send only the four passages that look relevant. 4k tokens instead of 500k, refreshed by re-indexing one document. That’s RAG. Why that breaks: “looks relevant” is doing an enormous amount of work, and each of its failures has a standard patch.

flowchart LR
    subgraph index [Offline: build the index]
        C[Corpus] --> CH[Chunk] --> E[Embed] --> IX[(Vector index)]
    end
    Q[Query] --> R[Hybrid retrieve<br/>embeddings + keyword]
    IX --> R
    R --> RR[Re-rank<br/>~100 → ~5] --> P[Prompt assembly] --> M[LLM] --> A[Answer]

The left half runs offline and can be updated incrementally; the right half runs per query. Retrieval survives huge context windows because that right-hand path sends ~4k relevant tokens instead of the whole corpus on every turn.

The shape of the problem shifts depending on what you’re retrieving over. Code retrieval cares about symbol graphs and call edges, not just embedding similarity. Conversation retrieval cares about recency. Legal retrieval cares about exact phrasing. There isn’t one “RAG”; there’s a family of pipelines that share the same shape.

The seam: agents complicate the picture

The interesting alternative to classical RAG isn’t long context — it’s agentic retrieval. Instead of one retrieval call before generation, the model uses tools to search, read files, follow links, and decide what to look at next. On the handbook question it might search “parental leave”, find the benefits index, open the leave section, notice a 2025 amendment link, and follow that too — none of which a single pre-generation lookup would have done. The corpus stays external, but the model drives retrieval instead of a fixed pipeline.

This is closer to how a human researcher works: you don’t pre-fetch the whole library, and you don’t read the whole book — you skim, follow citations, and stop when you have enough. Except that the human knows when to stop, and an agent doesn’t reliably: the failure mode moves from “the retriever picked the wrong chunk” to “the agent searched five times, found nothing, and answered anyway.” Whether agentic retrieval largely replaces classical RAG, or whether the two stay layered (agent on top, search underneath), isn’t settled — and the teams running these systems at scale don’t generally publish the comparisons that would settle it.

You started with RAG = retriever + generator + prompt assembly. What did the walk-through add? — + a ranking stage, and + provenance. The first is where most of the accuracy gets won or lost; the second is what lets you tell a grounded answer from the confident invented one you started with. And it’s the pair of them, not the size of the context window, that decides whether the handbook question gets answered correctly — which is why a bigger window didn’t retire any of this.

Going deeper