← all articles

How to chunk documents for RAG without wrecking retrieval

The chunk is the unit of retrieval, not an implementation detail

Most RAG post-mortems start the same way. The model gives a confident wrong answer, someone checks the retrieved context, and the actual fact was sitting in the source document the whole time. It just never made it into a chunk that got retrieved. The generation step gets blamed, but the failure happened three steps earlier, at chunking.

This matters because of how retrieval actually works. You embed a chunk, store the vector, and later compare it against a query embedding using cosine similarity or dot product. The chunk boundary decides what gets compressed into that single vector. If the boundary lands in the middle of a fact, the vector represents neither the fact nor its context well, and it won’t score high against the query that’s looking for it. You can have a great embedding model and a great reranker and still lose because the raw material handed to them was cut wrong.

Fixed-size chunking is the default, and the default is often wrong

The simplest approach, and the one most tutorials ship with, is splitting text every N characters or tokens, often 500 to 1000 tokens with no regard for sentence or paragraph boundaries. It’s fast, deterministic, and easy to implement in a few lines with a library like LangChain’s CharacterTextSplitter.

The problem is that fixed-size splitting doesn’t know what a sentence is. It will happily cut a chunk mid-sentence, mid-table-row, or right after a heading with none of the section body attached. When that half-sentence gets embedded, the vector represents a fragment that means something different from the full sentence, sometimes nothing coherent at all. Retrieval on fragments is noisier than retrieval on complete thoughts, because the embedding model was trained on natural language, not on arbitrary substrings.

Fixed-size chunking isn’t wrong in every case. For dense, uniform text like log files or transcripts without much structure, it’s a reasonable default because there’s no structure to respect anyway. For anything with headings, lists, tables, or code, it’s usually the wrong first choice.

Recursive splitting respects structure before it respects size

Recursive character splitting, the approach behind LangChain’s RecursiveCharacterTextSplitter and similar tools elsewhere, tries a list of separators in order: split on double newlines first (paragraphs), then single newlines, then sentences, then words, only falling back to a hard character cut if nothing else fits under the size limit. The result is chunks that end at natural boundaries whenever the document allows it, while still respecting a maximum size so you don’t end up with a 4000-token chunk because one paragraph happened to be that long.

This is a meaningfully better default than pure fixed-size splitting for most prose documents: docs, articles, policy text, support tickets. It doesn’t require understanding the content, just the document’s punctuation and whitespace, which makes it cheap to run at scale. The tradeoff is that it still doesn’t know what the content means. Two paragraphs that are tightly related conceptually but separated by a section break will end up in different chunks with no signal connecting them.

Semantic chunking follows meaning instead of punctuation

Semantic chunking embeds individual sentences, then groups adjacent sentences whose embeddings are similar into the same chunk, and starts a new chunk when the similarity between consecutive sentences drops below a threshold. The idea is that a topic shift in the text produces a jump in embedding space, and that jump is a better chunk boundary than an arbitrary token count.

It works, but it costs more to run because you’re embedding at the sentence level before you ever get to chunk-level embeddings, and it’s slower on large corpora. It also inherits the sensitivity of whatever threshold you pick: too tight and you get chunks of one or two sentences that lack context on their own, too loose and you’re back to chunks that span unrelated topics. This is a technique worth reaching for on long, meandering documents like research papers or long-form articles where topic boundaries don’t line up with paragraph breaks. It’s overkill for short, already-structured content like FAQ entries or API docs, where the structure is already doing the work semantic chunking is trying to approximate.

The embedding model has a say too

Chunk size isn’t just a document property, it’s constrained by the embedding model you’re using. Every embedding model has a maximum input length, and most degrade well before they hit that hard limit. Cram a chunk that’s too long into the model and the resulting vector becomes an average over everything in it, diluting the specific fact you wanted retrievable. This is the mechanical reason “just make chunks bigger so you never split a fact” doesn’t actually solve the problem: a long chunk gets you one vector representing several ideas at once, and a query about any single one of those ideas competes against noise from the others in the same vector.

Check your embedding model’s documented input limit and treat it as a ceiling, not a target. If you’re chunking at 2000 tokens because you want to avoid mid-fact splits, you’re better off addressing the split problem directly, through better boundary detection or context attachment, than making chunks big enough to blur past it.

Small chunks lose context, big chunks lose precision

This is the core tension and there’s no chunk size that avoids it entirely. Small chunks, in the 100 to 300 token range, retrieve precisely: the vector represents one specific claim, so a query about that claim scores it highly. But a small chunk handed to the LLM in isolation often lacks the surrounding context needed to interpret it correctly. “The limit is 15%” retrieved without knowing which limit, on what, under what conditions, is close to useless.

Larger chunks, 500 to 1000 tokens, carry more context but retrieve less precisely, because the vector is now an average over more content and a narrow query has to compete with everything else packed into that chunk. Neither end is free. Picking a chunk size is picking which failure mode you’d rather debug, and that choice should be informed by how your queries actually look. Short, factual lookup queries favor smaller chunks. Queries that ask for explanation or reasoning across a passage favor larger ones.

Tables, code blocks, and lists need their own rules

Generic text splitters, fixed or recursive, tend to butcher tables and code. A table split across two chunks loses the header row in the second chunk, so a row of numbers gets embedded with no label telling you what those numbers are. Code blocks split mid-function lose the signature or the closing brace, and the fragment doesn’t parse as anything meaningful to either the embedding model or a human skimming retrieved context.

The fix is format-aware preprocessing before your general splitter runs: parse tables and keep header plus rows together as one unit regardless of size, and split code by function or class boundary using the language’s actual syntax rather than a token count. This adds engineering work per content type, but it directly addresses one of the most common silent failure modes in document-heavy RAG systems, where the missing fact was in a table the pipeline had already mangled before embedding.

Attach context instead of hoping the chunk carries it

A chunk pulled from the middle of a document usually doesn’t know what document it came from. “Applicants must submit form 27B within 30 days” is a different fact depending on which policy document it’s from and which section that requirement belongs to, but a plain chunk doesn’t carry that back with it.

The fix that generalizes well is prepending a short piece of context to each chunk before embedding it, typically the document title, the section heading, and sometimes a one-line summary generated for that specific chunk. This means the embedding for the chunk reflects both its content and its place in the document, and it means the text handed to the LLM at generation time is self-contained instead of a floating fragment. It costs a bit more in embedding tokens and, if you’re generating per-chunk summaries, in LLM calls during ingestion. That cost is one-time per document, not per query, so it’s usually worth paying.

A chunking strategy you can actually test

Chunking decisions are easy to make in the abstract and hard to validate without real feedback. The only way to know if a chunking strategy is working is to build a small set of real queries with known correct source passages, run retrieval, and check whether the right chunk shows up in the top results. This doesn’t require a big benchmark suite, even twenty representative queries pulled from real usage will surface obvious failures, like a fact that never gets retrieved because it’s permanently split across a chunk boundary.

Rebuild this check whenever you change chunk size, overlap, or splitting method. Chunking strategy is not a decision you make once at the start of a project and forget. It’s a parameter you tune the same way you’d tune any other part of the pipeline, against evidence from your own queries and your own documents, not against a default someone else picked for a different corpus.

If you want more breakdowns like this on RAG pipelines, LLM tooling, and the AI coding assistants people are actually shipping with, come find us at AI Tool Gazette.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →