← all articles

When a reranker earns the latency it costs you

Every RAG tutorial eventually tells you to bolt a reranker onto your retrieval step. Fewer of them tell you what that costs. If you’ve shipped a pipeline and watched your p95 latency jump after adding one, you already know the tradeoff is real. The question worth answering isn’t “should I use a reranker,” it’s “does my pipeline’s failure mode actually match what a reranker fixes.” Most of the time nobody checks.

What a reranker actually does differently from your retriever

Your first-stage retriever, whether it’s a vector index or a hybrid BM25 plus embedding setup, works because document embeddings are precomputed. You embed your corpus once, offline. At query time you embed the query and compare it against those stored vectors with cosine similarity or dot product. That comparison is cheap because the heavy lifting (running text through a model) already happened before the user asked anything.

A reranker, almost always a cross-encoder, can’t do that trick. A cross-encoder takes the query and a candidate document together, as a single input, and runs them jointly through the model to produce a relevance score. Because the query isn’t known ahead of time, you can’t precompute anything. Every candidate you want scored means a full forward pass through the model, at query time, on the critical path.

That’s the entire reason reranking is more accurate and also the entire reason it’s slower. Joint attention between query and document tokens lets the model catch relationships a separate-embedding comparison misses. It also means the cost scales with how many candidates you feed it, not with corpus size.

Where the latency actually goes

A typical pipeline looks like: retrieve top 50 to 100 candidates from your vector store, hand them to the reranker, keep the top 3 to 10, pass those to the LLM. The reranking step has two latency sources stacked on top of each other. First, compute: N candidates means N (or N batched) forward passes through a transformer, and that scales close to linearly with N. Second, if you’re calling a hosted reranking endpoint rather than running the model in-process, you’ve added a network round trip on top of the compute time, plus whatever queuing happens on the provider’s side.

If you’re scoring 100 candidates through a hosted API, you’re paying for 100 candidates’ worth of compute even though your context window only reads the top 5. That mismatch is where a lot of wasted latency hides. Nobody budgets for it because the retrieval step “already gave us a ranked list,” so the reranker feels like a formality instead of a second real inference call.

When the extra hop is worth it

Reranking earns its cost when your first-stage retrieval is genuinely bad at ordering, not just bad at recall. A few situations where that’s true:

Your corpus has a lot of near-duplicate or similarly-phrased chunks. Technical documentation, legal text, and API references are full of passages that use the same vocabulary to describe different things. Embedding similarity flattens a lot of that nuance because it’s comparing whole-chunk vectors, not token-level interaction. A cross-encoder reads query and document together and can pick up on the distinction.

Your embedding model is generic and untuned for your domain. If you’re using an off-the-shelf embedding model on niche content, first-stage retrieval quality is going to be mediocre no matter how you tune the vector store. Reranking gives you a second, differently-shaped model to catch what the first one missed.

Your downstream context window is small relative to your candidate pool. If your synthesis step only reads the top 3 to 5 chunks, then what lands in positions 1 through 5 is the entire game. A reranker that moves the actually-relevant chunk from position 12 to position 2 changes what the model sees. If you’re stuffing the top 30 chunks into a huge context window regardless, the reranker’s reordering matters a lot less because the LLM reads the misranked chunk anyway.

Your chunking produces a lot of borderline-relevant fragments. Multi-hop questions, long documents split into overlapping windows, and FAQ-style content with repeated phrasing all create candidate pools where the top 20 by embedding score are close in relevance. That’s exactly the regime where joint query-document attention adds signal that vector similarity can’t.

When it’s just tax

The flip side matters just as much. Reranking doesn’t help, and just adds a round trip, in a few common cases:

Your corpus is small and your embeddings already separate cleanly. If you’ve got a few thousand well-structured documents and the right chunk consistently lands in the top 2 or 3 of first-stage retrieval, there’s nothing for a reranker to fix. You’re paying compute and latency to reorder a list that was already correct.

You’re already over-fetching into a large context window. If your synthesis step reads the top 20 or 30 chunks regardless of order, reordering the top 100 down to that set doesn’t change much about what the model sees, because most of what the reranker would promote was already included.

Your product is latency-sensitive chat, not batch synthesis. In a live chat UX, users feel every added round trip directly. If reranking buys you a marginal precision gain but the retrieval was already good enough to produce a correct answer, you’ve traded a UX cost for a quality gain nobody notices.

Your actual problem is upstream of ranking. A reranker can only reorder the candidates it’s given. If your chunking strategy is splitting sentences mid-thought, or your indexing pipeline is missing whole sections of source documents, no amount of reordering fixes that. You’re polishing a list that’s missing the answer.

The mechanical way to decide, not vibes

Don’t guess. Look at where the actually-relevant chunk lands in your current top-k, across a sample of real queries you can label by hand. If the right chunk is consistently sitting in position 1 or 2, reranking has nothing to do. If it’s scattered across positions 5 through 30, that’s precisely the range a reranker is built to fix, since it’s the range your retriever ranks inconsistently but a joint-attention model can often sort out.

Candidate count is the other lever, and it’s one people forget they control. Since cross-encoder cost scales with how many candidates you score, cutting your pre-rerank pool from 100 down to 20 through better first-stage filtering or metadata pre-filters cuts reranker latency by roughly the same factor, without giving up the reordering quality you actually need. Reranking 20 well-chosen candidates is a different latency budget than reranking 100 mediocre ones.

A cheaper middle ground

If the latency budget is tight but you still see real ranking errors in your evaluation, you don’t have to choose between “always rerank everything” and “never rerank.” A lighter cross-encoder model scores faster than the largest one you can find, and for a lot of domains the accuracy gap is small enough not to matter. You can also gate reranking behind a cheap heuristic, only invoking it when the top-k scores from first-stage retrieval are close together, since a wide score margin usually means the retriever was already confident and correct. And if your traffic has repeat queries, caching reranked results for identical or near-identical queries removes the cost entirely on the second hit.

The honest tradeoff

A reranker isn’t free quality. It’s quality you’re buying with a second inference call on the critical path, and whether that purchase is worth it depends entirely on whether misranking is actually your bottleneck. Measure where your relevant chunks land today before you add one. If they’re already near the top, you’re paying latency for nothing. If they’re buried in the tail of your candidate list, that’s the exact problem a reranker was built to solve, and the round trip is the price of fixing it.

If you’re working through tradeoffs like this one across your RAG stack, coding assistants, or the rest of your AI tooling, you can find more breakdowns like this on the AI Tool Gazette home page.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →