Why your retrieval system gives confidently wrong answers
I spent about six weeks improving prompts for a retrieval system that had never handed the model the right document.
The system answers questions over my own operational documentation: setup guides, pricing, carrier notes, a few years of support replies. When it started producing answers that were wrong in specific, plausible ways, I did what everyone does. I rewrote the system prompt. I added examples. I swapped the model twice.
Then I logged what the retriever was returning and the whole problem became visible in an afternoon.
Print the passages first
Not a count of them. Not a summary. The actual text of every retrieved chunk, in rank order, with scores, next to the question that produced them.
Read them yourself and answer one question: is the information needed to answer this anywhere in that text?
For roughly two thirds of the questions that were failing on mine, it was not. The correct chunk was not in the top ten. The model had been asked to produce an answer from material that did not contain one, and it obliged, because that is what these things do when you do not tell them otherwise.
Six weeks of prompt work could not have fixed a single one of those.
The number that caps your entire system
Take fifty questions people actually asked. Not questions you wrote by skimming your own docs, which is a different and much easier test. Real ones, out of a support inbox or a chat log.
For each, open the corpus and note by hand which document contains the answer. That labelling is an hour of tedious work and there is no way around it.
Then run each question through retrieval and check whether that document appears in what came back.
The resulting percentage is a hard ceiling on your finished system. No amount of prompting, no larger model, no clever output parsing can lift an answer above the material it was given. Most teams I have compared notes with have never calculated this number once, and several were surprised by how low it was.
Chunking beats embedding model choice, and it is not close
I changed embedding models three times. Each swap moved my results by an amount I could barely distinguish from noise.
Changing how documents were split moved everything.
The tutorial default is to cut text every N characters. It is two lines of code and it is the worst decision in most pipelines. Here is what it did to me: a setup guide with numbered steps got split so that step four landed in a chunk on its own, without the “before you begin” paragraph three steps above it that said which hardware had to be powered first. The chunk retrieved perfectly for questions about step four. It was also, read on its own, instructions for doing the thing in the wrong order.
Tables are worse. Split a pricing table between the header row and the body and you have produced a chunk of numbers with no column labels. Mine had 200 and 500 sitting in a column with a few dollar figures, and the model had no way to know which were gigabytes and which were prices. It guessed, fluently.
What fixed it:
- split on structures the document already has: headings, section breaks, list boundaries, whole tables. Never split a table.
- prepend the heading path to the chunk text. Document title plus section name costs about thirty tokens and turns a floating fragment into something locatable.
- overlap the edges by a sentence or two, which catches definitions that sit just before a boundary.
There is a ten minute test for all of this. Pull twenty chunks at random and read them without the source document open. If you cannot tell what a chunk is about, the embedding that has to represent it cannot either.
Vector search cannot find a serial number
This one is structural, not a bug, and it caught me out because it is the opposite of what embeddings are good at.
An embedding places text in a space of meanings, which is exactly what you want when a user asks in their own words and your document uses different ones. It is exactly wrong when a user pastes a string they want matched literally.
I searched my notes for a specific modem model number and got five documents about modems in general. The page with that exact number in it did not appear. Same failure for error codes, invoice references, version strings, surnames. The embedding smears an identifier into a neighbourhood of similar looking identifiers, which is precisely what the user did not want.
Run a plain keyword search alongside the vector search and merge the result sets. If you are already on Postgres, full text search is sitting there and costs you nothing extra to operate.
My position, which some people will argue with: if your users ever type identifiers, keyword search is not an enhancement you get to defer. Without it a slice of your queries can never succeed, no matter what else you build on top.
Reranking is still the best value change available
Retrieve thirty candidates instead of five, score them with a reranking model, keep the best four.
The reason it works: the first search has to be fast across the whole corpus, so it uses a cheap approximation of relevance. The reranker only sees thirty passages, so it can afford to read them properly. Broad and cheap, then narrow and careful.
Cost is a few hundred milliseconds and one small model call. It did more for my answer quality than every hour of prompt work combined, and it is about fifteen lines of code. I deferred it for months because it sounded like added complexity.
A stale index produces no error at all
The first genuinely bad answer my system gave quoted a data cap of 500GB. The real figure had been 200GB since March. I had rewritten that page and never reindexed it.
Nothing in the pipeline knew. Retrieval succeeded, ranking looked healthy, the answer was specific and well written and four months out of date. The only reason I caught it is that a human who knew the real number happened to read it.
Three things that help:
- store a hash or modified timestamp per source document alongside its chunks, and run a job that compares and reindexes what moved.
- handle deletions. Most indexing scripts only add. Mine had chunks from pages that had not existed for a year, because nothing ever removed them.
- put the document date into the chunk text. Then the model can say the figure is from March and the reader gets to decide whether that is good enough. An answer carrying a date is a different object from one without.
When two of your documents disagree
I had two versions of the same pricing information in the corpus. One current, one from a page I had forgotten about. Both retrieved, both went into the prompt.
The model picked one and said nothing about the other. No flag, no hedge, same confident register it uses when the sources agree.
Which one wins is close to arbitrary. In my case the stale document ranked higher because it was longer and repeated the query terms more often, so being wrong made it look more relevant.
Deduplicate near identical chunks and keep a source date you can sort on. Then add the instruction: if retrieved passages contradict each other, say so and quote both instead of choosing. It will do that when asked. It will never do it unprompted.
Teach it to refuse, then test the refusal
Tell the model to answer only from the supplied passages, and to say plainly that it does not know when the answer is not there. Most people write some version of this and consider it handled.
It is not handled until it is measured, and measuring it needs an evaluation set almost nobody builds.
My first eval set was fifty questions and every one of them was answerable. That is what happens when you write test questions by reading your own documents. The system scored well on it. It scored well on the only thing I was measuring while the failure that actually mattered, inventing an answer out of nothing, sat completely outside the test.
I was optimising a number that could not move when the real problem got worse.
Rebuild the set so a third of the questions have no answer in your corpus. Plausible questions, right subject area, genuinely uncovered. The correct response to each is an admission of ignorance, and anything else is a scored failure. That gives you a fabrication rate, which is the number worth watching.
The order I check in now
- Is the answer in the corpus at all? Sometimes the document was never ingested and everything downstream is theatre.
- Is the correct passage in the top ten? Measure it against hand labelled questions.
- Does each retrieved chunk make sense read cold?
- Does the query contain a literal string that needs keyword matching?
- Is the index current, and does anything ever remove deleted material?
- Do any two retrieved passages contradict each other?
Only after all six do I look at the prompt or think about the model.
Two things I have not solved. Questions needing two documents joined together still beat me: asking what a given plan costs and whether it includes a particular feature requires a hop, and one search over the question text finds neither passage cleanly. And automatic grading of long answers works until the judging model shares the blind spots of the model it is grading, which mine did. So I still read a sample by hand every week, which takes twenty minutes and has caught things no metric did. More of what I have measured on tools like these, with the bills attached, is here.