← all articles

Hybrid search: why pure vector retrieval keeps missing

The support ticket that started this

A few months into running a RAG pipeline over a product’s internal docs, I watched it fail on a query like “error E4021 on startup.” The vector index had the right document sitting three rows down in the corpus, word for word, error code and all. It just never got retrieved. The embedding model had encoded “E4021” into a vector space built for semantic meaning, and semantic meaning doesn’t know what to do with an alphanumeric code that shows up exactly once in the entire dataset. Cosine similarity pulled back documents about “startup errors” in general instead of the one document that actually contained the string E4021.

That’s the moment most people building on top of vector databases learn the same lesson: pure vector retrieval is not a universal search solution. It’s a specific tool with a specific blind spot, and that blind spot happens to cover exactly the kind of query that shows up constantly in real usage: IDs, SKUs, error codes, acronyms, exact proper nouns, version numbers.

What dense vectors are actually good at

An embedding model takes text and maps it to a point in high-dimensional space, trained so that semantically similar text ends up close together. Ask “how do I reset my password” and it can retrieve a document titled “account recovery steps” even though the two phrases share almost no words. That’s the entire value proposition, and it’s real. Paraphrase, synonym, and conceptual matching are things keyword search structurally cannot do.

The tradeoff is that embeddings compress meaning, and compression throws away precision. A model with 768 or 1536 dimensions is representing an entire sentence as a single point. Rare tokens, exact strings, and numbers get smoothed into the surrounding semantic neighborhood rather than preserved as literal matches. The model has never seen “E4021” enough times in training to give it a distinct, stable representation, so it gets treated more like noise than a a specific search anchor.

What keyword search is actually good at

BM25, the scoring function behind most keyword search including Elasticsearch’s default relevance scoring, works completely differently. It counts term frequency (how often a query word appears in a document), weights that by inverse document frequency (how rare that word is across the whole corpus), and normalizes for document length so a long document doesn’t win just by containing more words. There’s no learned representation, no training data, no semantic understanding at all. It’s counting and weighting.

That’s exactly why it nails exact matches. “E4021” is a rare term, so IDF weights it heavily, and any document containing that literal string jumps to the top. The same logic applies to product SKUs, function names in code documentation, legal citations, and person names. If the user’s query and the target document share the exact token, BM25 finds it reliably, with no training required and no risk of the token getting diluted into a vector.

The tradeoff is the mirror image of vector search. BM25 has no idea that “car” and “automobile” mean the same thing. A query about “reset password” will not match a document that only says “regain account access” even though a human reading both would call them the same request.

Why this isn’t a fixable gap in either method alone

It’s tempting to think you can just pick the better embedding model and close the gap, or fine-tune on your domain vocabulary until rare terms get their own stable vectors. That helps at the margins, but it doesn’t change the underlying mechanism. Dense retrieval is fundamentally a similarity operation in continuous space. Continuous space is bad at representing “this exact discrete token must appear.” You’d need the embedding to essentially memorize every rare identifier in your corpus as a distinct direction in vector space, which defeats the purpose of using embeddings for generalization in the first place.

Same story in reverse for keyword search: no amount of synonym dictionaries or stemming rules gives BM25 real semantic understanding. You can hand-add synonym lists, but that’s manual curation that never keeps pace with how users actually phrase things.

These are two different retrieval mechanisms solving two different problems. Hybrid search exists because production queries are a mix of both problem types, often within the same query.

How hybrid search actually combines them

The mechanics are simpler than the marketing around them suggests. You run the same query through both retrieval paths, a BM25 (or similar lexical) search and a dense vector search, in parallel. Each path returns its own ranked list of candidate documents with its own scores. The hard part is combining two ranked lists that use completely different scoring scales, since a BM25 score and a cosine similarity score aren’t comparable numbers.

Two approaches show up most often in real systems:

Weighted score combination. Normalize both score types into a common range and combine them with a weight, something like final_score = alpha * vector_score + (1 - alpha) * bm25_score. This works but is fragile, because the right alpha depends on your query distribution and tends to need retuning as your corpus or query patterns shift.

Reciprocal rank fusion (RRF). Instead of combining raw scores, RRF combines rank positions. A document gets a score of 1 / (k + rank) from each retrieval path it appears in, and those get summed. A document that ranks highly in both lists wins even if the two scoring systems are on totally different scales, because rank position is comparable across methods in a way raw scores aren’t. This is why RRF shows up as the default fusion method in most hybrid search implementations, including the ones built into Elasticsearch, Weaviate, and Qdrant. It sidesteps the score normalization problem entirely.

A lot of production pipelines add a third stage after fusion: a cross-encoder reranker that takes the top candidates from the fused list and scores query-document pairs jointly instead of independently. This catches cases where both retrieval paths agree on a mediocre document and both miss a better one, because the reranker actually reads the query and candidate together rather than comparing precomputed representations.

What this costs you in practice

None of this is free. Running two retrieval paths means maintaining two indexes, an inverted index for the lexical side and an ANN index (HNSW, IVF, or similar) for the vector side, which roughly doubles your indexing infrastructure and the engineering surface area you’re responsible for. Fusion adds a step to your query path, and if you add a reranker on top, that’s another model call, another few hundred milliseconds of latency, and another API bill if you’re using a hosted reranking model rather than running one locally.

The question worth asking before you build this is whether your query distribution actually needs it. If your corpus is genuinely narrative, prose-heavy content with no identifiers, codes, or exact-match terms that matter, pure vector search might be enough and hybrid search is added complexity for a problem you don’t have. If your users search for order numbers, error codes, model names, or anything else that behaves like an exact token rather than a concept, you will hit the E4021 problem, and no amount of embedding model upgrades fixes it on its own.

The takeaway for anyone building RAG

Vector search and keyword search aren’t competing approaches where one is the upgrade path from the other. They fail on different query types for structural reasons baked into how each method represents text. If you’re shipping a RAG pipeline that has to handle both “how do I fix a slow checkout” and “error E4021,” you need both retrieval paths, a fusion step that doesn’t assume one scoring scale, and a clear-eyed view of the latency and infrastructure cost that comes with running two indexes instead of one.

Want more breakdowns of how RAG pipelines actually work under the hood, without the vendor spin? Head back to the AI Tool Gazette home page for the rest of our explainers.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →