Advanced RAG

Course Content

Advanced RAG

3 sections · 38 lessons

What are multi-vector retrieval techniques, and why are they effective?


What you need to know

The problem with one vector per chunk

An embedding compresses a chunk into one point, typically 768–3,072 numbers. If the chunk has one topic, that point is a good summary. If it has several, the point sits somewhere between them. Rare but decisive tokens — a plan price, an error code, a SKU — have little influence on the average. Multi-vector retrieval keeps the separate aspects separate.

Family 1: several representations of the same content

For each chunk, store extra vectors, each with the same parent_id:

RepresentationWhat it catches
The chunk textQueries that use the document's own wording
An LLM summaryBroad "what is this about" queries
Hypothetical questionsInformal, question-shaped queries
Section title plus pathNavigation-style queries
Table rows or key-value factsPrecise lookups

At query time, search all vectors, map hits to parent_id, deduplicate, and return each parent once. LangChain's MultiVectorRetriever (now in langchain-classic) implements this pattern; it is also easy to build with any store that supports metadata.

Family 2: token-level late interaction

ColBERT stores one vector per token and scores with MaxSim — each query token finds its best-matching document token, and the best matches are added up. ColPali does the same with image patches of a page. This keeps word-level detail without running a cross-encoder at query time. Several vector databases now store multiple vectors per point and score with MaxSim natively.

Multi-representation

  • 3–6 vectors per chunk
  • Standard vector index works
  • Needs LLM calls at ingestion
  • Good for varied query styles

Token-level (ColBERT-style)

  • One vector per token: hundreds per chunk
  • Needs MaxSim support or a reranking step
  • No LLM calls; a special embedding model
  • Good for exact terms and dense detail

Controlling the cost

Index size and memory grow with the number of vectors — a few times for multi-representation, far more for token-level. The standard answer is a cascade:

  1. Cheap first stage — single-vector ANN plus BM25 over the whole corpus, top 100–200.
  2. Multi-vector scoring — score only those candidates with MaxSim or the extra representations.
  3. Rerank — optionally a cross-encoder on the top 20–30.

Research such as MUVERA (2024) goes further, converting multi-vector sets into single fixed-size vectors so a normal ANN index can do the first stage.

A real-life example

An Indian telecom's help centre has one page per plan. A typical page covers price and validity, daily data, calls and SMS, OTT subscriptions, international roaming and cancellation rules — six topics in 700 tokens, embedded as two chunks.

A customer asks: "₹599 wale plan mein international roaming hai kya?" (Does the ₹599 plan include international roaming?). The chunk containing the answer is mostly about OTT apps and calling, with one roaming line. Its single vector matches "OTT" questions well and roaming questions poorly; a general roaming FAQ ranks higher and gives a generic, wrong answer.

The team adds, per chunk, a summary vector and one vector for each extracted fact line ("₹599 plan: international roaming not included; add-on packs available"). The fact-line vector matches the question closely and points back to the plan chunk, which ranks first. Index size grows about four times, which on 12,000 articles is still small. On 400 real plan questions, answers about specific plan features become far more accurate, while broad questions such as "which plans have Netflix?" are unchanged.

Follow-up questions to expect

  • "How do you merge scores when several vectors of one parent match?" — Take the maximum score per parent (the best-matching representation wins), or fuse per-representation result lists with RRF.
  • "Is hybrid search a form of multi-vector retrieval?" — Loosely: BM25 is a second, sparse representation of the same chunk. Fusing it with dense search gives similar benefits for exact terms.
  • "When is it not worth it?" — Short, single-topic chunks gain little; one vector already represents them well.