Comparison · 5 minute read
Reranker Comparison: When a Second Pass Earns Its Latency
A reranker reorders retrieval results to raise precision, and it cannot recover a passage retrieval never returned. Compare on precision gain against your corpus, latency and cost added per query, how many candidates each can handle, and whether hybrid retrieval would help more for less.
A reranker raises precision on what retrieval found and cannot recover what it missed. This guide covers when the second pass earns its latency, drawing on FISTA Solutions' AI enablement retrieval work.
What should the comparison cover?
Six dimensions, measured on your corpus.
| Dimension | What to measure | Why it matters |
|---|---|---|
| Precision gain | Top-result accuracy before and after | The reason to add it |
| Latency added | Per query, at the tail | Interactive budgets |
| Candidate capacity | How many it can score | Limits the improvement |
| Domain fit | Gain on your content | General benchmarks mislead |
| Cost per query | Compute or API | Scales with volume |
| Deployment | Hosted or self-hosted | Portability and residency |
Why does it only affect precision?
Because it reorders a set it did not choose.
Retrieval decides which candidates exist; the reranker decides their order. A passage absent from the candidate set is absent from the output regardless of how good the reranker is.
That means recall problems must be fixed in retrieval — better embeddings, better chunking, hybrid search — before reranking is worth adding. Teams that add a reranker to fix a recall problem see no improvement. See RAG quality checklist.
How many candidates should you pass?
Enough that the reranker has room to work, bounded by latency.
Retrieving a few dozen candidates and reranking to the top handful is a common shape. Retrieving five and reranking to three gives the reranker almost nothing to do.
The trade is latency: more candidates means more scoring work. Test several candidate counts against your precision measurement and find where the gain flattens.
What does latency cost you?
Time on every query, before the model even starts generating.
Reranking runs after retrieval and before generation, so it is pure added delay from the user's perspective. On an interactive interface that goes directly into time to first token.
Measure it at your candidate count and concurrency, not in isolation. A reranker that is fast on ten candidates may not be on fifty. See AI performance tuning checklist.
Why does domain fit matter?
Because relevance judgement is domain-specific, as it is for embeddings.
A reranker trained on general web content may judge relevance differently from what your users mean in a technical, legal, or clinical corpus. The gain on your content is the only figure that matters.
Measure with your own query and passage pairs. The ranking of rerankers frequently reorders on specialised corpora, as it does for embedding models. See embedding model comparison.
Would hybrid retrieval help more?
Frequently, and for less latency.
Combining keyword and vector retrieval improves both recall and precision on business corpora full of exact terms — product codes, names, identifiers. That is a retrieval-layer fix addressing the more fundamental problem.
Try hybrid before reranking. If hybrid closes the gap, you have improved recall and precision without adding a per-query model call. See vector database comparison.
Hosted or self-hosted?
Rerankers are small enough that self-hosting is practical.
That removes per-query cost, keeps content within your boundary, and freezes the version. The operational burden is modest compared with serving a generative model.
For data residency requirements this is, like embedding, one of the easier components to bring in-house. See open weight vs hosted models for enterprise.
How do you run your own comparison?
Take your query and passage test set, measure top-result accuracy with retrieval alone, then with each candidate reranker at several candidate counts.
Record latency at the tail alongside the precision gain. If the gain is small and the latency is noticeable, the answer is no.
What does switching cost later?
Low. A reranker is a stateless scoring step, so replacing it does not touch stored data or require re-embedding.
That makes it one of the easiest components to change, which is an argument for measuring rather than agonising.
What do people get wrong here?
Adding one to fix a recall problem. Too few candidates to rerank. Latency unmeasured. Choosing on general benchmarks. And skipping hybrid retrieval, which frequently helps more.
Does the generation model make this unnecessary?
Partly. A capable model given ten passages can identify the relevant one, which reduces the cost of imperfect ordering.
But passing more passages costs input tokens on every call and dilutes attention, so precise retrieval remains worth having. The trade is reranking latency against input token cost, and it depends on your volume. See the context window arms race.
Which should you choose?
Fix recall first with better retrieval and hybrid search. Add a reranker only if measurement on your corpus shows a precision gain that justifies the latency. It is a cheap component to change, so decide by measuring rather than by reasoning.
What should you do first?
Measure top-result accuracy with your current retrieval. If the right passage is usually first, a reranker adds latency for little gain.
How FISTA Solutions helps
FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: recall addressed in the retrieval layer before reranking is considered, with precision gain and tail latency measured on the client's own corpus, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.
To run this comparison against your own workload, message FISTA on WhatsApp, or read how to improve RAG accuracy.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What does a reranker actually do?
Scores each retrieved candidate against the query with a model that sees both together, then reorders. That joint scoring is more accurate than comparing separate embeddings, at more computational cost.
02Can it fix poor retrieval?
No. It can only reorder what retrieval returned. If the right passage was not in the candidate set, no reranking recovers it, which is why recall must be addressed first.
03How many candidates should you retrieve?
More than you will use — commonly several times the final count — so the reranker has room to improve the ordering. Too few candidates and there is nothing to reorder.
04What does it cost?
Latency on every query plus per-query compute or API cost. For interactive interfaces the latency is the binding constraint more often than the cost.
05When do you not need one?
When retrieval already returns the right passage in the top position most of the time, or when hybrid retrieval would close the gap more cheaply. Measure before adding.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.