RAG: why retrieval didn't die when context windows got huge
Long context windows were supposed to kill retrieval-augmented generation. They didn't. Here's why the bottleneck moved instead of disappearing.
On this page
The picture version
Six pictures for a reader who has never heard the term. The prose below fills in the seams the pictures skip.
1 · The problem
Fluent, instant, and about somebody else’s company.
2 · First obvious fix
Train it on the handbook. Now it answers — and you still can’t check it.
3 · Second obvious fix
Paste in everything. Pay for it, every single message.
4 · The fix
Fetch four pages before answering, not the whole library.
5 · The seam
The librarian never reads the pages it hands you.
6 · Keep this card
The whole idea on one index card.
Why it exists
You ask the company chatbot “how much parental leave do I get?” and it answers immediately, in fluent HR prose, with a number. The number is wrong. It has never read your employee handbook. It read a great many other companies’ documents during training, and what came out is a plausible policy rather than yours. That handbook is the running example for this post.
That’s the wall apps built on top of an LLM ran into once they moved past demos: the model didn’t know your data. It knew a slice of the open web up to its training cutoff, and that was it. Your codebase, your customer’s tickets, your internal wiki, last week’s Slack — invisible.
The obvious fix was “just put the handbook in the prompt.” But the amount of text a model can read at once — its context window — was a few thousand words in 2022, and the handbook plus every other internal doc was a few million. So people did the next thing: cut the documents into passages, turn each passage into a list of numbers that captures roughly what it’s about, store those, and when a question arrives fetch only the passages that look relevant. Stuff those into the prompt. That’s Retrieval-Augmented Generation, and it became the standard answer to “how do I do AI on my own data?”
Then context windows grew, fast: 100k tokens in mid-2023, and a million with Gemini 1.5 in early 2024. And a recurring take appeared: RAG is dead, just paste the whole corpus in.
It didn’t happen. Retrieval didn’t go away, and the interesting question is why.
Why it matters now
If you’re building anything that talks to an LLM about non-public information — a coding agent, a support bot, a search-over-docs feature, a notes app — you’ll almost certainly end up with some retrieval step, even if you arrive at it by accident and call it something else. Understanding why the bottleneck moved instead of disappearing is the difference between a system that works at 100 documents and one that works at 100 million.
The short answer
RAG = retriever + generator + prompt assembly
Picture to keep: a librarian who pulls four pages for you before you ask your question, rather than handing you the whole library — or expecting you to have memorized it. Except that this librarian never reads the pages: it matches on shape and similarity, so the wrong four pages arrive silently, and the model answers from them just as fluently.
A retriever picks a small number of relevant chunks from a big corpus, an LLM generates an answer conditioned on them, and a thin layer in between decides what actually goes into the prompt. The reason huge context windows didn’t kill it: putting more tokens in the prompt costs money and latency on every turn, and even when you can afford it, the model gets worse at using them.
How it works
Take the parental-leave question and try to fix it the obvious ways.
Naive attempt: fine-tune the model on your handbook. Teach it the facts directly. Why it breaks: you have to redo it whenever HR edits the policy, and — the killer — the answer comes back with no provenance. The model can’t show you the paragraph it got “16 weeks” from, so you’re back to not knowing whether it’s real. Further training can teach facts; what it doesn’t give you is an inspectable link back to the source text, which is the thing you actually needed here. (A fine-tuned model can certainly print something citation-shaped. Whether that citation points at a real paragraph is exactly the question you were trying to settle.)
Fix: put the handbook in the prompt. Now the answer is grounded in text the model can quote, and updating the policy means updating a document. This is the right shape. Why it breaks at scale: three separate ways.
- You pay per token, every turn. The model doesn’t charge you for the window; it charges you for the tokens you actually send. A 1M-token window is a capacity, not a free lunch. Stuffing 500k tokens of internal docs into every request means paying to process 500k tokens on every message. Prompt caching helps when requests share the same opening stretch of text; the more the relevant slice changes per query — which is the whole point of search — the less of the prompt it can amortize.
- Models get worse at long contexts than the marketing implies. This is the “lost in the middle” / “context rot” finding: as the relevant fact moves from the start or end of a long prompt toward the middle, accuracy drops, sometimes sharply. Liu et al. (2023) measured this directly, and the qualitative pattern has held up across later evaluations. How large the effect is on any particular current model is harder to state: vendors publish their own retrieval-from-long-context scores under their own conditions, and those aren’t comparable across labs. The practical read: don’t assume that putting the expense policy next to the leave policy is free for the answer about leave.
- Corpora aren’t static. HR edits the handbook. New tickets arrive. The codebase gets a commit. An index can be updated incrementally; a monolithic prompt can’t.
Fix: retrieve first, then prompt. Chunk the corpus, embed each chunk, and at query time send only the four passages that look relevant. 4k tokens instead of 500k, refreshed by re-indexing one document. That’s RAG. Why that breaks: “looks relevant” is doing an enormous amount of work, and each of its failures has a standard patch.
- Embeddings miss exact strings. Ask about error code
E_PARENTAL_0032and semantic similarity shrugs — it has no notion that this token must appear verbatim. Fix: hybrid retrieval, combining semantic search (embeddings) with keyword search (BM25). Embeddings catch paraphrase; keywords catch identifiers. - The top 100 hits are still mostly noise. Vector search is fast because each chunk was embedded once, in isolation, before your query existed. Fix: re-ranking — a smaller, slower cross-encoder that scores each (query, chunk) pair and cuts the top ~100 down to the top ~5. It’s often the cheapest large quality win in a retrieval pipeline, because it fixes the ranking without touching the index.
- The answer straddles a chunk boundary. Split the handbook every 500 tokens and “16 weeks” ends up in one chunk while “parental leave” sits in the previous one; neither retrieves. Fix: better chunking — respecting document structure, overlapping windows, carrying section headers into each chunk. Chunking is an unglamorous knob that decides whether the right passage was ever retrievable in the first place; Anthropic’s contextual retrieval work reports measured retrieval-failure reductions from fixing exactly this.
- The user can’t check the answer. Fix: prompt assembly that carries provenance — which doc, which section — so the model can cite and the human can click.
flowchart LR
subgraph index [Offline: build the index]
C[Corpus] --> CH[Chunk] --> E[Embed] --> IX[(Vector index)]
end
Q[Query] --> R[Hybrid retrieve<br/>embeddings + keyword]
IX --> R
R --> RR[Re-rank<br/>~100 → ~5] --> P[Prompt assembly] --> M[LLM] --> A[Answer]
The left half runs offline and can be updated incrementally; the right half runs per query. Retrieval survives huge context windows because that right-hand path sends ~4k relevant tokens instead of the whole corpus on every turn.
The shape of the problem shifts depending on what you’re retrieving over. Code retrieval cares about symbol graphs and call edges, not just embedding similarity. Conversation retrieval cares about recency. Legal retrieval cares about exact phrasing. There isn’t one “RAG”; there’s a family of pipelines that share the same shape.
The seam: agents complicate the picture
The interesting alternative to classical RAG isn’t long context — it’s agentic retrieval. Instead of one retrieval call before generation, the model uses tools to search, read files, follow links, and decide what to look at next. On the handbook question it might search “parental leave”, find the benefits index, open the leave section, notice a 2025 amendment link, and follow that too — none of which a single pre-generation lookup would have done. The corpus stays external, but the model drives retrieval instead of a fixed pipeline.
This is closer to how a human researcher works: you don’t pre-fetch the whole library, and you don’t read the whole book — you skim, follow citations, and stop when you have enough. Except that the human knows when to stop, and an agent doesn’t reliably: the failure mode moves from “the retriever picked the wrong chunk” to “the agent searched five times, found nothing, and answered anyway.” Whether agentic retrieval largely replaces classical RAG, or whether the two stay layered (agent on top, search underneath), isn’t settled — and the teams running these systems at scale don’t generally publish the comparisons that would settle it.
You started with RAG = retriever + generator + prompt assembly. What did the walk-through add? — + a ranking stage, and + provenance. The first is where most of the accuracy gets won or lost; the second is what lets you tell a grounded answer from the confident invented one you started with. And it’s the pair of them, not the size of the context window, that decides whether the handbook question gets answered correctly — which is why a bigger window didn’t retire any of this.
Famous related terms
- Embeddings —
embedding = text → vector— the substrate that makes semantic search possible. See embeddings. - Vector database —
vector DB ≈ index optimized for nearest-neighbor in high dimensions— the storage layer. Note that it doesn’t have to be a separate product: Postgres (via pgvector) and Elasticsearch both do vector search, which is often enough. - Re-ranker —
re-ranker = cross-encoder + (query, doc) → score— the step that turns “100 plausibly relevant chunks” into “5 actually relevant ones.” - Lost in the middle —
lost in the middle = LLM recall sags for facts placed mid-prompt vs. at the ends— the empirical finding that LLMs underuse the middle of long contexts. The exact magnitude is model-specific and moves around with each release.
Going deeper
- Lewis et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (2020) — the primary source, and the answer to “did RAG originally mean what it means now?” (no: the retriever was trained jointly with the generator, not bolted on at prompt time).
- Anthropic, Contextual Retrieval (2024) — answers “how much does chunking actually cost me, and what fixes it?”, with measured retrieval-failure rates rather than folklore.
- Liu et al., Lost in the Middle (2023) — the rabbit hole, for “does where I put a fact in the prompt change whether the model uses it?” Frontier models have improved on the answer; they haven’t erased it.