← all articles

What a bigger context window does not solve

context-window retrieval llm prompt-caching

I took the input cap off my documentation assistant the week whole files started fitting in a single call. After years of squeezing everything into a few thousand tokens, it felt free.

It was not free. Cost per answer roughly tripled and I found out at the end of the month, from the invoice, which is the worst available place to find out anything.

The answers did get better, and I want to be fair about that before I start arguing. Every time the window grows, somebody announces that retrieval is finished, and I keep having to be the person who says it is not.

What a large window genuinely fixed

You stop cutting documents into pieces that should never have been cut.

My clearest case is a supplier quote. Two pages of line items, then an appendix of exclusions that changes what half those line items mean. Split it into paragraph sized fragments and retrieve one line item, and you get a price with no conditions on it. Retrieve the appendix and you get conditions attached to nothing.

Now the whole quote goes in and the model reads it the way a person would, in order, with the appendix in view.

Runbooks are the other one. Mine open by defining which host is which, then say things like “roll the second one back first” eleven paragraphs down. A fragment carrying that sentence is worse than useless. It reads like a complete instruction, and the one thing that makes it safe to follow is missing.

Cross references inside a document resolve again. That is a real gain and I have no interest in minimising it.

You rent the window by the call

Every token you put in is billed again the next time somebody asks something. A large window used without a budget is a standing charge nobody approved, and that is the whole of my cost point here, because I have written up where agent bills actually come from separately.

The middle of a long input is the worst place to put something

Attention over a long sequence is uneven. Material near the beginning and material near the end get used more reliably than material sitting in the middle, and that comes out of how the mechanism works rather than out of a bug in one vendor’s release.

I tested it on my own corpus in the least sophisticated way available. Longest carrier document I have, the same fact planted at the top, in the middle, and near the end, twenty questions asked from each position. The middle placement lost. Not dramatically. Enough that I would not put a billing answer on it.

The practical consequence is the one people miss. Filling the window can make an answer harder to reach. You add thirty thousand tokens hoping to cover every case, the passage you actually needed lands somewhere in the middle of the pile, and you have buried it with the material you added to help.

Most long prompts I see are laid out worst first: conversation history, then tool schemas, then documents, then the question somewhere inside all of it. Every individual piece is justifiable. The layout is indefensible.

Put the material most likely to matter at the ends. Which is a ranking decision, which is the thing you were told you could stop doing.

There is a happy accident here worth taking. If you cache the stable prefix of a prompt, the unchanging parts have to sit first and the question ends up last anyway. Cost and attention want the same layout, which almost never happens.

A window is not a store

It resets. Every call starts from nothing.

Anything the system needs to know next week has to be written down somewhere and read back, which is exactly the machinery people expect a large window to replace.

I hit this in a small way. I had told the assistant, mid session, that one carrier needs the APN set by hand on new SIMs. Useful. Gone by the next morning, because there was nowhere for it to have gone.

The usual patch is to summarise the history and carry the summary forward. It helps. It also discards the detail you will want later, silently, and you find out which detail that was when somebody asks for it.

I ended up keeping a short file per customer thread holding what we agreed. That is a database with bad ergonomics, and writing it was the moment I stopped believing the window had replaced anything.

Do the arithmetic on your own corpus

Pick a window size. Any of them. Now go and measure your document set.

Mine is a few hundred files plus two years of support mail, and it does not come close to fitting. Mine is also small. Anyone with a legal folder, a decade of internal wiki and a serious support history has orders of magnitude more than I do.

The two numbers grow at different speeds, which is the part that settles it for me. Window sizes step up when a vendor ships, a few times a year. Corpora grow whenever a human does anything, which is continuously. I have never seen one shrink.

So you are choosing what goes into the call. You cannot opt out of choosing, because the alternative is a call that will not fit.

That choice is retrieval. It makes no difference whether you implement it as embeddings, as a keyword filter, or as a loop over a directory in alphabetical order until you run out of room. The last one counts too. It is retrieval with no ranking function and nobody measuring it.

The version most teams actually ship is pasting the whole repository, or the whole shared drive, into the prompt. That works on a small one and then stops, and it stops abruptly. One day the call fits, the next day it does not, and now you are writing selection logic under time pressure with somebody waiting.

There is one honest exception. If your entire corpus fits with headroom and you are content paying for all of it on every question, skip retrieval. One product manual. A staff handbook. A single contract. That category is real and it is not where most people are.

The unit changed, the job did not

Three years ago I retrieved paragraphs, because a paragraph was all that fit alongside everything else the call needed. Most of my effort went into splitting rules, which I was bad at and disliked.

Now I retrieve documents. The search still has to pick the right file. Picking the wrong one still ruins the answer. The difference is that what arrives is intact.

That is the accurate description of what the window did. It raised the unit of retrieval from a paragraph to a document. Genuine improvement, considerably smaller claim, and it survives contact with a real corpus.

Those years on splitting rules were not wasted, they were spent one layer too low. The same judgement now goes into deciding what counts as a document at all. For me that meant taking one enormous notes file and breaking it into eleven titled ones.

How my index is built now

One embedding per file, generated from the title plus a short summary I compute once and store, instead of one embedding per thousand characters.

Two or three whole documents per call, ordered so the file the search is most confident about sits last, right up against the question.

A hard cap on total input. If what I want to send goes over the budget I set for a question, the weakest document gets dropped instead of the call getting bigger. That cap is the design decision almost everyone skips, and without it your cost per question drifts upward every time somebody adds a long file, with nothing anywhere to tell you.

Four or five of my files are still too long to send whole, so the chunker is still in the codebase doing about a tenth of its old job. I did not get to delete it. I got to demote it.

Here is the position I get argued with about. For any corpus I cannot read myself in an afternoon, I want a retrieval layer even when everything would fit. The reason is debugging. When an answer comes out wrong I want a short list of what the system decided to look at. With the whole corpus in the window there is no list, and the only move left is to reword the question and hope.

What I still cannot do

The middle of the window finding bothers me, because my evidence for it is thin. One document, one planted fact, twenty questions per position. Enough to make me order my inputs carefully, nowhere near enough to publish a number, and it almost certainly moves around by model and by how the document is written. My ordering rule is a hunch with a small sample behind it, and I would rather say so.

Questions that span two documents are still beyond me. Document level retrieval improved single document questions a lot and did nothing for the case where the answer needs the rate card and a carrier note held side by side. I get both when the search happens to be lucky. I have no reliable way of making it lucky.

And the cap I took off in the first paragraph took me a month to notice, which is the real risk in all of this. The window does not fail loudly. It gets quietly more expensive while the answers get better, and that is the hardest kind of regression to argue anybody out of. The full setup and what each change did to the monthly bill is here.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →