How to chunk documents for retrieval
Most bad RAG answers I get asked about are not model problems. The right passage is in the index, sliced in half at a paragraph break, and neither half scores high enough to reach the prompt.
This is for anyone building question answering over their own documents (a help centre, contracts, runbooks, a folder of PDFs) who can run a Python script. You need a chunker and a test set to score it against.
By the end you’ll have a chunking script, 20 to 30 test questions, a recall@5 number you can compare between settings, and a frozen config. Budget an hour or two.
What you need
- Python 3.10 or newer, then
pip install langchain-text-splitters tiktoken pypdf openai numpy - 30 to 200 representative documents in a folder. Use messy real ones.
- an embedding model. I use OpenAI’s text-embedding-3-small below. It was $0.02 per million tokens at the time of writing, so a 1M token test corpus costs about two cents. A local model works the same way, see how to choose an embedding model, and running a model on your own hardware if documents can’t leave your network.
- somewhere to keep vectors. A numpy array in memory is fine for now. Qdrant or Weaviate come later (weaviate vs qdrant for rag).
- 20 to 30 real questions your users ask
Step by step
1. Extract clean text first
Convert every source into plain text or markdown, one file per document, in a corpus/ folder. Extraction is where most chunking problems start.
# extract.py
import re
from pathlib import Path
from pypdf import PdfReader
Path("corpus").mkdir(exist_ok=True)
for pdf in Path("raw").glob("*.pdf"):
pages = (p.extract_text() or "" for p in PdfReader(pdf).pages)
text = "\n\n".join(pages)
text = re.sub(r"(?<!\n)\n(?!\n)", " ", text) # pdf line wraps are not paragraphs
Path("corpus", f"{pdf.stem}.md").write_text(text, encoding="utf-8")
print(pdf.name, len(text))
pypdf returns text without headings. If your PDFs have real structure, pymupdf4llm outputs markdown with headings, which step 4 uses. For web pages, trafilatura does the same job, and the proxyscraping.org blog covers fetching them.
Expected output: one line per file with a character count.
If it breaks: a count of 0 or a few dozen characters means a scanned PDF. Run ocrmypdf on it first.
2. Write your test questions before you chunk
Create questions.json with 20 to 30 real questions. Each one gets a must field: a short exact phrase, 5 to 10 words, copied from the passage that answers it. Good sources are support tickets, site search logs, and the queries you log (see what to log in an AI application).
[
{"q": "how do i rotate an api key", "must": "select Regenerate next to the key"},
{"q": "what is the refund window", "must": "within 14 days of purchase"}
]
A retrieved chunk only counts as a hit if it contains the whole phrase, so a boundary that slices the answer shows up as a miss.
Expected output: a file with 20 or more entries.
If it breaks: if you can’t find a passage for a question, that’s a content gap, not a chunking problem. Drop the question.
3. Pick a size in tokens
Set size in tokens, not characters. I start at 400 tokens with 60 tokens of overlap and only move once I have a score to compare. An embedding model squashes a chunk into one vector, so a chunk covering three topics gets a blurry vector that matches none well.
OpenAI’s embedding models accept about 8,000 tokens per input (embeddings guide), but that’s a ceiling, not advice. On the generation side, Lost in the Middle (Liu et al., 2023) found models use information at the start and end of a long context better than the middle. That’s my reason for sending five tight chunks instead of two huge ones.
Expected output: two constants at the top of your script, CHUNK = 400 and OVERLAP = 60.
If it breaks: if your documents are mostly short FAQ entries, 400 is too big. Try 200.
4. Split on structure first, then on size
Split on markdown headings first so a chunk rarely straddles two topics. The size splitter then only steps in for sections over 400 tokens, trying paragraphs, then lines, then words. Tiny leftovers under 80 tokens merge back into the previous chunk of the same section.
# chunk.py
import json
from pathlib import Path
import tiktoken
from langchain_text_splitters import MarkdownHeaderTextSplitter, RecursiveCharacterTextSplitter
CHUNK, OVERLAP, MIN_TOKENS = 400, 60, 80
enc = tiktoken.get_encoding("cl100k_base")
by_header = MarkdownHeaderTextSplitter(
headers_to_split_on=[("#", "h1"), ("##", "h2"), ("###", "h3")],
strip_headers=False,
)
by_size = RecursiveCharacterTextSplitter.from_tiktoken_encoder(
encoding_name="cl100k_base", chunk_size=CHUNK, chunk_overlap=OVERLAP
)
chunks = []
for path in sorted(Path("corpus").glob("*.md")):
for sec in by_header.split_text(path.read_text(encoding="utf-8")):
section = " > ".join(sec.metadata[k] for k in ("h1", "h2", "h3") if k in sec.metadata)
for piece in by_size.split_text(sec.page_content):
small = len(enc.encode(piece)) < MIN_TOKENS
prev = chunks[-1] if chunks else None
if small and prev and prev["doc"] == path.name and prev["section"] == section:
prev["text"] += "\n\n" + piece
else:
chunks.append({"doc": path.name, "section": section, "text": piece})
Path("chunks.json").write_text(json.dumps(chunks, indent=1), encoding="utf-8")
sizes = sorted(len(enc.encode(c["text"])) for c in chunks)
print(f"{len(chunks)} chunks, median {sizes[len(sizes) // 2]}, max {sizes[-1]}")
Expected output: a line like 1240 chunks, median 210, max 470. A median well under 400 is normal, and a max a little over 400 comes from merged leftovers.
If it breaks: a median under 60 means extraction left one line per paragraph. Go back to step 1.
5. Put the heading path in front of every chunk
At embed time, prepend doc > section to each chunk. A chunk that says “select Regenerate” means little alone. Prefixed with api-keys.md > Rotating keys, the embedding knows what it’s about. It’s free and usually helps; step 8 measures it.
The bigger version is Anthropic’s contextual retrieval: an LLM writes a short blurb placing each chunk in its document, and you prepend that. Their write-up reports 35% fewer failed top-20 retrievals from contextual embeddings alone, and 49% with contextual BM25 added. It costs one LLM call per chunk, so try the free prefix first.
Expected output: the embedder text starts with the path, as in the next script.
If it breaks: documents without headings give an empty section. The prefix falls back to the filename, which still helps.
6. Embed and score with recall@5
# eval.py
import json
import numpy as np
from openai import OpenAI
client = OpenAI()
chunks = json.load(open("chunks.json", encoding="utf-8"))
questions = json.load(open("questions.json", encoding="utf-8"))
def embed(texts):
vecs = []
for i in range(0, len(texts), 256):
r = client.embeddings.create(model="text-embedding-3-small", input=texts[i:i + 256])
vecs += [d.embedding for d in r.data]
m = np.array(vecs, dtype="float32")
return m / np.linalg.norm(m, axis=1, keepdims=True)
docs = embed([f'{c["doc"]} > {c["section"]}\n{c["text"]}' for c in chunks])
qs = embed([q["q"] for q in questions])
top = np.argsort(-(qs @ docs.T), axis=1)[:, :5]
hits = sum(any(q["must"] in chunks[i]["text"] for i in row) for q, row in zip(questions, top))
print(f"recall@5: {hits}/{len(questions)}")
Expected output: something like recall@5: 21/26.
If it breaks: AuthenticationError means OPENAI_API_KEY isn’t set in your shell. A rate limit error means lower the batch size from 256. A perfect score on the first run usually means your must phrases are too generic.
7. Read every miss
Add this to the end of eval.py:
for q, row in zip(questions, top):
if not any(q["must"] in chunks[i]["text"] for i in row):
homes = [i for i, c in enumerate(chunks) if q["must"] in c["text"]]
print(q["q"], "| phrase lives in chunks:", homes or "NONE", "| top1:", chunks[row[0]]["section"])
Every miss lands in one of three buckets. NONE means the phrase was split across a boundary or lost in extraction, which is a chunking bug. A chunk id that just missed the top 5 is a ranking problem, so try the context blurb or a different embedding model. And if the question wording shares nothing with the document wording, that’s a retrieval problem, not a chunking one.
Expected output: one line per miss, sortable into those buckets.
If it breaks: if every miss says NONE but the text looks fine, whitespace differs. Normalise both sides with " ".join(s.split()).
8. Change one variable, then freeze
Re-run steps 4 and 6 with one change at a time: chunk size 200, 400, 800; overlap 0, 60, 120; heading prefix on or off. Keep a small table of scores. Chroma’s team published a chunking evaluation in 2024 comparing chunkers and sizes. Read it for how they scored things, and as a reason to test overlap instead of assuming it helps.
I skip semantic chunking (splitting where embedding similarity drops) by default. I haven’t run a big comparison, so that’s a lean, not a finding. If headings plus the size splitter get you past 90% on your set, I wouldn’t pay for extra embedding calls.
Then freeze. Write the settings to chunk_config.json with a version string and store that version on every chunk.
Expected output: a table where one setting wins clearly, or all sit within a question or two of each other. With 25 questions, one question is 4 points, so a one-question gap is noise.
If it breaks: if every setting ties, chunking isn’t your bottleneck. Look at the embedding model, or add BM25 and a reranker.
Common pitfalls
- Sizing in characters. The splitter counts characters unless you use
from_tiktoken_encoder. An English token is roughly four characters, but non-Latin scripts tokenise very differently, so mixed-language corpora give uneven chunks. - Tuning on the three queries you tried by hand. A score on 25 questions beats a feeling.
- Cutting tables, code blocks and numbered procedures in half. A step list split at step 4 answers “how do I” with half a procedure.
- No document id or content hash on each chunk. When a page changes, or someone asks you to remove it, you can’t find its chunks.
- Indexing personal data with no way to delete it. A vector store holds copies of your text, so removal means deleting chunks as well as the source file. In Singapore the PDPA may apply. This is not legal advice, and the privacywire blog covers data protection more broadly.
Scaling this
At 10x, a few thousand documents, the numpy array still holds up. Add a content hash per document so you only re-chunk what changed, and rerun the eval on every config change.
At 100x, say 100,000 documents of 2,000 tokens each, that’s 200M tokens. Embedding them at $0.02 per million is $4, so embedding isn’t the cost. LLM-generated context is: one call per chunk, across millions of chunks. Read how to read an LLM provider pricing page properly before you start. A real vector database earns its place here, and so do hybrid search and a reranker.
At 1000x, extraction and re-indexing dominate, with PDF extraction and OCR slowest. Run chunking as a queue of idempotent jobs keyed by document hash. Changing the chunker means re-embedding everything, so build into a new collection, run your eval, then swap. That’s why step 8 stamps a version on every chunk.
Where to go next
- how to choose an embedding model: the two interact, so re-run your eval when you switch
- weaviate vs qdrant for rag: where the vectors live once you outgrow numpy
- what to log in an AI application: the retrieval logs that feed your next test set
- the full article index
Written by Xavier Fok
disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-09-22.