Glossary ┬╖ 5 minute read
What Is Query Expansion? Improving Retrieval Recall Explained
Query expansion rewrites or augments a query before retrieval to improve what it finds: adding synonyms, generating several formulations, or embedding a hypothetical answer. It raises recall on short or vaguely worded queries and risks drifting away from what the user actually asked.
Retrieval frequently fails not because the index is poor but because the query is short, vague, or uses different vocabulary from the documents. Query expansion addresses that at the query side rather than the index side, which is cheaper to iterate on and easy to overdo. This explainer covers the techniques and their trade-offs. It complements how to improve rag accuracy and what is retrieval augmentation, and reflects FISTA Solutions' approach in AI agents delivery.
What problem does it solve?
The vocabulary gap. A user asks how to stop a payment going out; the document explains cancelling a standing instruction. They share almost no terms, and a short query gives retrieval very little to work with.
Expansion bridges this without asking users to phrase things better, which they will not do.
| Technique | Cost | Best for |
|---|---|---|
| Synonym addition | Minimal | Known terminology gaps |
| Multi-query generation | One generation, several retrievals | General improvement |
| Hypothetical answer | One generation | Question-document mismatch |
| Conversational rewriting | One generation | Multi-turn sessions |
| Term weighting | Minimal | Emphasising key terms |
| Decomposition | One generation | Compound questions |
What is multi-query generation?
Producing several reformulations of the question, retrieving for each, and merging results. One phrasing may match the documentation's language, another the user's, another an intermediate.
It is the most consistently effective technique because it hedges across phrasings rather than betting on a single rewrite. The cost is several retrieval operations plus one generation, which is acceptable for most workloads and noticeable at very high volume.
What is hypothetical answer embedding?
Generating a plausible answer to the question and embedding that rather than the question. The insight is that a generated answer looks structurally like the document that genuinely answers it тАФ declarative, using domain vocabulary тАФ while the question does not.
It works well and it surprises people, because the generated answer may be factually wrong and still retrieve the right document. Correctness of the hypothetical is not the point; resemblance is.
Why does conversational rewriting matter?
Because follow-ups are incomplete. "What about the second one" or "does that apply in France" retrieve nothing useful in isolation, and multi-turn retrieval simply does not work without rewriting them into self-contained queries using the conversation history.
This is the least optional of the techniques. Any system supporting follow-up questions needs it, and systems that lack it produce baffling behaviour on the second turn.
What is the risk?
Drift. An expanded query can retrieve documents relevant to the expansion rather than to the original question, and the resulting answer is confidently about something slightly different from what was asked.
Hypothetical answers carry this risk most, since a generated answer that goes in the wrong direction retrieves documents supporting the wrong direction. Precision should be measured alongside recall rather than assuming expansion is free. See what is an evaluation rubric.
How should it be measured?
Recall and precision together, on a labelled query set, comparing each technique against no expansion. Expansion that raises recall by four points and drops precision by six is a loss, and only measuring both reveals it.
Latency matters too: each generation adds time before retrieval even begins, which users experience directly.
What should you do first?
Add conversational rewriting if you support follow-up questions, because it fixes a class of failure rather than improving an average. Then test multi-query generation against your labelled set before anything more elaborate.
How does it interact with filtering?
Carefully, because an expanded query can imply a different scope from the original. If a rewrite introduces a jurisdiction or product that the user did not mention, and filters are derived from the query, the system may search the wrong subset entirely.
The safe arrangement derives filters from the original query and the authenticated context, never from the expansion. Expansion should widen what matches within a scope, not change the scope itself.
Should the expansion be visible?
Not usually to the user, and always in the logs. Showing a rewritten query mid-conversation is confusing; recording it is essential, because when an answer is off-target the expansion is frequently the cause and it is invisible without a trace.
That log line is one of the cheapest debugging aids in a retrieval system, and it is regularly missing from implementations where the expansion was added as a quick improvement rather than as a designed step.
When should it be skipped?
When queries are already specific and well-formed, which is common in expert-facing tools where users know the terminology. Expansion adds latency and drift risk for no gain there, and applying it uniformly across a system that serves both expert and general users is worse than applying it selectively.
How does it affect cost and latency?
Directly. A generation step before retrieval adds latency the user experiences before anything appears, and multi-query approaches multiply retrieval operations. On high-volume interactive systems that cost is significant enough to justify applying expansion selectively тАФ on queries that look short or ambiguous тАФ rather than universally.
How FISTA Solutions helps
FISTA Solutions implements conversational rewriting as a baseline for multi-turn systems, tests expansion techniques against labelled query sets measuring precision alongside recall, and accounts for the added latency before adopting generation-based expansion, through AI agents, AI enablement, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.
To close the gap between how users ask and how your documents are written, message FISTA on WhatsApp, or read how to improve rag accuracy.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What problem does expansion solve?
The gap between how users ask and how documents are written. A short query using everyday words may share almost no vocabulary with the formal document that answers it, and expansion bridges that without requiring users to phrase things better.
02What is multi-query generation?
Producing several reformulations of the question, retrieving for each, and merging the results. It is the most consistently effective technique because it covers multiple phrasings rather than betting on one, at the cost of several retrieval operations.
03What is hypothetical answer embedding?
Generating a plausible answer to the question and embedding that instead of the question. The generated answer resembles the document that actually answers it more closely than the question does, which improves matching even though the generated content may be factually wrong.
04Why does conversational rewriting matter?
Because follow-up questions are incomplete. What about the second one retrieves nothing useful on its own, and rewriting it into a self-contained query using the conversation history is what makes multi-turn retrieval work at all.
05What is the risk?
Drift. An expanded query can retrieve documents relevant to the expansion rather than to what the user asked, which produces confidently off-target answers. Precision should be measured alongside recall rather than assuming expansion is free.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. WeтАЩll map the fastest credible path from intent to verified production.