← all articles

When a bigger window replaces your retrieval layer

The pitch nobody asked for

Every few months someone ships a model with a bigger context window and a chunk of the internet declares RAG dead. It isn’t dead. But the calculus around when you actually need a retrieval pipeline versus when you can just paste the whole document set into the prompt has genuinely shifted, and most teams haven’t sat down and recalculated it. This is that recalculation, based on how both approaches actually work, not on a chart from a launch blog post.

What RAG actually does under the hood

Retrieval augmented generation is, mechanically, a search problem wearing an AI costume. You take your source documents, split them into chunks (usually a few hundred tokens each, with some overlap so you don’t slice a sentence in half), and run each chunk through an embedding model to get a vector. Those vectors go into an index, usually something doing approximate nearest neighbor search like HNSW, because exact nearest neighbor search over millions of vectors is too slow to be useful.

At query time, you embed the user’s question, pull back the top-k closest chunks (k is often something small like 5 to 20), sometimes run a reranker over that shortlist to reorder by relevance, and stuff the survivors into the prompt. The model never sees your corpus. It sees whatever your retrieval step decided was relevant.

That last sentence is the whole story. RAG’s quality ceiling is retrieval quality, not model quality. If the chunk that answers the question never makes it into the top-k, the model can be as capable as you like and it still won’t find the answer, because it was never shown the answer. This is why RAG systems fail quietly on questions that don’t phrase themselves the way the source text does, and why teams end up bolting on keyword search (BM25) alongside vector search as a hybrid, because pure embedding similarity misses exact matches like product SKUs, error codes, or names that don’t carry much semantic weight but matter enormously to the user.

What a bigger context window actually does under the hood

A long context model skips retrieval entirely. You put the whole document, or several documents, directly into the prompt, and the model attends over all of it when generating a response. There’s no chunking decision to get wrong, no embedding model to pick, no index to keep in sync with your source files. The tradeoff is that “attend over all of it” is not free, and it is not uniform.

Self-attention’s cost grows with the length of the input, so processing a long prompt takes real prefill time before the model produces its first output token. This is why a call with a 150,000 token document attached to it noticeably takes longer to start responding than a call with a 2,000 token snippet, even if the actual answer is one sentence long. It’s also why cost scales with what you send: you’re billed on input tokens, and if you resend the same 150k token document on every turn of a conversation, you pay for it every single turn, not once. Providers have started offering prompt caching for repeated prefixes specifically because teams kept hitting this bill and asking why. Caching helps, but only if your workflow reuses the same fixed context across calls rather than assembling something new each time.

Where RAG still wins

If your corpus doesn’t fit in a context window at all, the conversation is over before it starts. A support knowledge base with 40,000 articles, a legal firm’s document archive going back fifteen years, a codebase with millions of lines, none of that fits into even a 1 million token window, and it never will just because windows keep growing, because corpora grow too. Retrieval is the only mechanism that scales to “search across more data than any single model call can hold.”

RAG also wins when your data changes constantly and you need answers reflecting the current state, not a snapshot. Updating a vector index with new or edited documents is a targeted, incremental operation. Rebuilding a giant static context every time a document changes is wasteful in a way that compounds fast if updates happen daily.

And RAG wins on cost at scale for narrow questions. If most user queries only need a handful of paragraphs to answer, paying to reprocess an entire document set on every call is money spent buying nothing, because the model was never going to use 95% of what you handed it.

Where long context wins

Long context wins on multi-hop reasoning, which is the case retrieval structurally struggles with. If answering a question requires connecting a fact in document A to a fact in document B, and neither document is semantically similar to the query on its own, your retriever may never surface both chunks together. The model, given the whole set at once, can make that connection because it isn’t relying on a similarity score to decide what’s worth reading.

It also wins when your working set is bounded and you actually want the model to see everything, not a filtered subset. A single contract, a pull request diff plus the files it touches, a handful of PDFs for a research summary, these are cases where the “documents” are small enough to fit whole, and where a retrieval step would only introduce a chance of dropping something the answer depends on. This is also, practically speaking, how a lot of AI coding assistants work day to day: reading a file in full rather than embedding it into chunks and hoping the relevant function ranks in the top five. Grepping for a symbol to find which file to open is retrieval in the loosest sense, but once the file is found, the whole thing gets read, not sliced.

The lost in the middle problem

Long context isn’t a free upgrade to “the model reads everything perfectly.” Researchers studying long-context recall have repeatedly documented that models are more reliable at using information near the start or end of a long prompt than information buried in the middle of it, a pattern generally referred to as “lost in the middle.” That means dumping 300,000 tokens of loosely relevant material into a prompt and hoping the model finds the one paragraph that matters is not a guaranteed win just because the tokens technically fit. Position and signal-to-noise inside the window still matter. A context window is capacity, not comprehension.

A hybrid that most teams end up shipping

In practice, a lot of production systems land somewhere between the two extremes rather than picking a side. Instead of fine-grained chunking and reranking, they do coarse retrieval, pull back whole files or whole documents based on metadata like file path, folder, or document type, and then let a big context window hold and reason over that coarser set. You get retrieval’s scalability for narrowing down which documents matter, without the failure mode of a chunker cutting a document into fragments that lose their surrounding context. It’s a reasonable middle ground when your corpus is too large to paste in whole but small enough that “which five documents” is an easier problem than “which five hundred-token slice of which document.”

So which one do you use

Ask what actually breaks first: if it’s the size of your corpus, you need retrieval, and no context window size fixes that. If it’s the quality of your retrieval, meaning the model keeps missing answers that exist in your data but never got surfaced, a bigger window that lets you skip retrieval for that specific bounded task will outperform a broken pipeline every time. Neither one is the default answer. The size of your data and the shape of your questions decide it, not which approach shipped a bigger number in its launch announcement this quarter.

For more breakdowns like this on AI tools, RAG pipelines, and coding assistants without the vendor spin, head back to the AI Tool Gazette home page.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →