What actually happens when your document is bigger than the context window
If you’ve built anything that touches real documents, you’ve hit this wall: the PDF is 400 pages, the transcript is three hours long, the codebase is 200 files, and the model’s context window is not big enough to hold all of it alongside your prompt and its answer. This isn’t an edge case. It’s the default state of working with long document LLM pipelines once you move past toy demos.
The fix isn’t one trick, it’s a set of tradeoffs. Here’s how they actually work, not how the marketing copy describes them.
What a context window actually limits
A context window is a token budget, not a page count. Every model has a maximum number of tokens it can hold across the system prompt, your input, and its output combined. Tokens aren’t words. Rough rule of thumb for English text: about 4 characters per token, so 100,000 tokens is somewhere around 75,000 words, not 100,000. PDFs with tables, code, or non-English text tokenize less efficiently than plain prose, so a “50 page PDF” can eat more tokens than you’d expect from the page count alone.
The budget also isn’t just the document. Your system prompt, conversation history, tool definitions, and the model’s own output all draw from the same pool. A model advertised with a large window rarely means you get all of it for your document. Leave headroom for the answer, or you’ll get a response that cuts off mid-sentence.
When the document doesn’t fit, you have three real options: make the window bigger, cut the document down before it gets to the model, or fetch only the relevant slice at query time. Most production systems end up doing some combination of the last two.
Option one: just use a longer context window
The obvious move is to pick a model with a bigger window and stuff the whole document in. This works, and for a lot of use cases it’s genuinely the right call, it’s simpler than building a retrieval pipeline and there’s no chunking logic to get wrong.
The catch is cost and latency scale with input size regardless of whether the model actually needs most of what you sent it. If you’re paying per token, and you are, sending a 300-page contract on every single question about it is expensive when the question only touches one clause. Latency also grows with input length because the model has to process every token before it produces the first output token. On a chat interface, that’s the difference between an instant reply and a multi-second wait before anything appears.
There’s also a well documented behavior where models are less reliable at pulling out details from the middle of a very long input than from the beginning or end. This isn’t a benchmark I ran, it’s a widely reported pattern in how attention behaves over long sequences, and it means “it fits in the window” is not the same guarantee as “the model will actually use all of it correctly.” Stuffing the whole document in front of the model is a starting point, not a solved problem, once the document gets long enough.
Option two: retrieval augmented generation (RAG)
RAG flips the problem around. Instead of sending the whole document, you split it into chunks, usually a few hundred to a couple thousand tokens each, embed each chunk into a vector, and store those vectors in a vector database. At query time, you embed the user’s question, find the chunks whose vectors are closest to it, and send only those chunks to the model along with the question.
This is the standard architecture behind most “chat with your documents” products, and it’s the right tool when the corpus is too large to ever fit in a context window at all, think a knowledge base with thousands of documents, not one long file. The cost per query stays roughly constant no matter how big the underlying corpus grows, because you’re only ever paying for the chunks you retrieve, not the whole library.
The tradeoff is retrieval quality becomes your bottleneck instead of context length. If the chunk that actually answers the question doesn’t get retrieved, because the chunking split it awkwardly or the embedding similarity missed it, the model never sees it and will either say it doesn’t know or, worse, guess. Chunk size and overlap matter more than people expect: chunk too small and you lose surrounding context that made a sentence meaningful, chunk too large and you dilute the vector so it’s less likely to match a specific query. Overlapping chunks (repeating the last 10-20% of one chunk at the start of the next) helps avoid cutting a key sentence exactly at a boundary, at the cost of some redundant storage and embedding calls.
RAG also adds real infrastructure: an embedding model, a vector store, and a re-embedding job every time the source documents change. That’s not free in engineering time even if the per-query API cost is low.
Option three: summarize the document down first
For cases where you need the model to reason over the whole document rather than answer a narrow lookup question, summarization based approaches often work better than retrieval. The two common patterns are map-reduce and hierarchical (recursive) summarization.
Map-reduce splits the document into chunks, summarizes each chunk independently (the “map” step), then feeds those summaries into another pass that combines them into a final summary or answer (the “reduce” step). It parallelizes well since each chunk’s summarization is independent, but you’re now making one LLM call per chunk plus a final combining call, so costs multiply with document length and you lose any information that didn’t survive the first-pass summary.
Hierarchical summarization does this in layers: summarize chunks, then summarize groups of those summaries, then summarize that, until the final layer fits in one context window. This handles genuinely huge documents better than a flat map-reduce because each layer’s input is bounded, but each layer of compression loses more detail. Ask a hierarchically-summarized pipeline for a specific number buried in the original text and it may not have survived to the top layer at all.
Both approaches trade completeness for cost and speed. They’re good for “give me the themes” or “what’s the overall sentiment” style tasks and bad for “what was the exact clause in section 4.2” style tasks, which is exactly the inverse of where RAG tends to be strong.
What most teams actually run in production
In practice, the systems I’ve seen ship combine these rather than picking one. A common pattern: RAG for retrieval of the relevant sections, then feed those retrieved chunks (not the whole document) into a long-context model so it has enough surrounding material to reason well, rather than a single isolated 300-token chunk with no context. Another common pattern for document review tools: run a first summarization pass to build a table of contents or section index, then use that index to decide which full sections to pull into context for the actual detailed question, instead of embedding the whole document into a vector store from the start.
None of this is free of tuning. Chunk size, overlap, number of retrieved chunks (top-k), and whether to re-rank retrieved chunks before sending them to the model are all knobs that need to be set based on your actual documents and query patterns, not copied from a tutorial. A pipeline tuned for legal contracts with long, dense sections behaves differently than one tuned for chat transcripts with short, choppy turns.
The practical decision
If your documents are small enough to fit in a single context window with room to spare, and you’re not running thousands of queries a day against them, just send the whole thing. It’s simpler and it avoids the failure modes of retrieval and summarization entirely.
If you’re querying a large, growing corpus where most questions only touch a small slice of it, build RAG, and budget real time for tuning chunking and retrieval, not just standing up the vector database.
If you need the model to reason across the entire document rather than answer a targeted lookup, summarization, ideally hierarchical, is usually a better fit than either of the above, with the understanding that you’re trading detail for coverage.
There’s no single correct architecture here. The question to ask is what kind of question your users are actually going to ask, because that determines whether you need everything, a slice, or a compressed version of everything.
If you want more breakdowns like this on what’s actually happening under the hood of the AI tools people are shipping with, check out the rest of AI Tool Gazette.