Glossary · 5 minute read
What Is Multi-Vector Retrieval? Multiple Embeddings Explained
Multi-vector retrieval represents a document with several embeddings instead of one: per passage, per token, or per anticipated question. It recovers detail that a single averaged vector loses, improves recall on long or multi-topic documents, and costs proportionally more index size and query work.
Retrieval quality problems on long documents usually have a specific cause: the document is represented by one vector that averages everything it contains, so no precise question matches it well. Multi-vector approaches address that directly, and they range from what most systems already do to techniques with substantial cost. This explainer covers the range. It complements what is late interaction retrieval and what is chunk overlap, and reflects FISTA Solutions' approach in AI agents delivery.
What does a single vector lose?
Specificity. An embedding is a fixed-length summary, and summarising a long multi-topic document into one produces something near the document's average meaning and near nothing in particular.
A twenty-page operations manual covering access, escalation, maintenance, and reporting yields a vector that is not close to a precise question about any of them. The answer is in the document and the document does not rank.
| Approach | Vectors per document | Recall gain | Cost |
|---|---|---|---|
| Single document vector | 1 | Baseline | Lowest |
| Passage or chunk vectors | Tens | Large | Proportional |
| Parent-child | Tens, plus parents | Large | Higher storage |
| Hypothetical questions | Tens more | Moderate | Ingestion generation |
| Token-level late interaction | Hundreds | Largest | Highest |
| Summary plus passages | Tens plus one | Moderate | Low |
What is the simplest form?
Chunking. Splitting a document into passages and embedding each one is already multi-vector retrieval, and nearly every production system does it. Framing it that way is useful because it reframes the question: not whether to use several vectors, but how many and of what kind.
Most of the available gain comes from this step. The more elaborate approaches refine it rather than replacing it.
What is late interaction retrieval?
Representing query and document as many token-level vectors and scoring by matching each query token against the document's tokens, then aggregating. Because matching happens at token granularity rather than between two summaries, it captures detail that single-vector similarity cannot.
The cost is real: storage grows by roughly the token count, and scoring is heavier. It is usually deployed as a reranker over a smaller candidate set rather than as the primary index. See what is late interaction retrieval.
What are hypothetical question embeddings?
Generating, at ingestion, the questions a passage would answer, and embedding those alongside the passage text. A user's question then matches a question rather than a declarative paragraph, which closes the gap between how people ask and how documents are written.
It works well where documentation is written formally and users ask casually, which describes most enterprise corpora. It costs a generation pass at ingestion and additional index entries.
What about summary vectors?
Embedding a generated summary of a document alongside its passages gives a representation that captures the whole while the passages capture the parts. It is cheap — one extra vector per document — and helps with questions about what a document is about, which passage-level retrieval answers poorly.
What does it all cost?
Index size, proportionally, plus the query work of searching more vectors and deduplicating results that come from the same source. Late interaction is the expensive end; hypothetical questions and summaries are modest additions.
The costs are predictable, which makes the decision straightforwardly measurable: does recall improve enough to justify the index growth on this corpus.
How should results be consolidated?
By deduplicating to the source document or passage before assembling context. Multiple vectors from one document will frequently all rank highly for a relevant query, and passing near-identical content several times wastes context budget without adding information.
The consolidation step also decides what the model actually sees — the matched chunk, its parent section, or the whole document — which is an independent choice worth making deliberately.
When is it not worth it?
On short, single-topic documents. A corpus of one-page FAQ entries gains almost nothing from elaborate multi-vector representation, because one vector already represents the content adequately. The gain scales with document length and topic diversity.
What should you do first?
Check whether your recall problem is concentrated in long documents. If precise questions fail on manuals and succeed on short notes, the diagnosis is representation rather than embedding quality, and passage-level retrieval with parent-child assembly will usually fix it at modest cost.
How does this interact with hybrid search?
Well, and independently. Multi-vector representation improves the dense side of a hybrid system; the sparse side is unaffected and continues to catch exact identifiers that no embedding handles. The two improvements compound rather than overlapping, which is why systems with good chunking and hybrid retrieval outperform those that invested heavily in only one.
The practical sequencing is hybrid retrieval first, because it is cheaper and fixes the most visible failures, then representation work on whatever recall gap remains.
What should you measure?
Recall at the retrieval depth you actually use, broken down by document length. A single aggregate recall figure hides the pattern that matters here, which is that short documents were always fine and long ones were not. Measuring separately tells you whether representation work is warranted and how much of the corpus it will help.
How does ingestion cost change?
Generating hypothetical questions or summaries adds a model call per chunk at ingestion, which is a one-off cost per document but a real one for large corpora. It also means ingestion becomes a pipeline with failure handling rather than a batch of embedding calls, and re-ingestion after a model change costs the same again.
How FISTA Solutions helps
FISTA Solutions diagnoses recall problems by document length and topic diversity before adding representation complexity, uses passage and parent-child retrieval as the default, adds hypothetical questions where user phrasing diverges from documentation, and deduplicates to source before assembling context, through AI agents, AI enablement, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.
To improve recall on long documents without over-engineering, message FISTA on WhatsApp, or read how to improve rag accuracy.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Why is one vector per document insufficient?
Because a single vector must average everything the document says. A twenty-page manual covering ten topics produces a vector near none of them specifically, so a precise question about one topic matches it poorly even though the answer is inside.
02What is the simplest multi-vector approach?
Embedding each passage or chunk separately, which nearly every production RAG system already does. Strictly this is multi-vector retrieval, and recognising it as such clarifies that the question is how many vectors per document, not whether to use several.
03What is late interaction retrieval?
Representing text as many token-level vectors and scoring by matching query tokens against document tokens individually before aggregating. It is considerably more accurate than single-vector similarity and considerably more expensive in both storage and computation, so it usually runs as a reranker over a smaller candidate set.
04What are hypothetical question embeddings?
Generating the questions a passage would answer and embedding those alongside the passage. It closes the gap between how users ask and how documents are written, at the cost of a generation step during ingestion and more vectors in the index.
05When is the extra cost justified?
When measured recall on real queries improves enough to matter, which depends on document length, topic diversity, and how differently users phrase questions. On short, single-topic documents the gain is often negligible and the cost is not.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.