Course Content
Advanced RAG
3 sections · 38 lessons
How does RAG Fusion improve recall over single-query retrieval?
What you need to know
Why one query misses
A dense retriever turns the question into one vector and returns its nearest neighbours. If the user and the document describe the same thing with different words, the two vectors can sit far apart. A document ranked 40th for the original wording may rank 3rd for a paraphrase. Keyword search (BM25) has the same problem in a sharper form: no shared word, no match.
The pipeline
- Generate — a small, fast model writes 3–5 alternative phrasings: synonyms, expanded acronyms, a more formal version, sometimes a translation.
- Retrieve in parallel — search once per query, including the original. Parallel calls keep the added latency close to one search.
- Fuse — combine the lists with RRF: each document scores the sum of
1 / (60 + rank)over the lists it appears in. - Rerank and trim — pass the fused top 30–50 to a cross-encoder and keep the top 5.
RRF uses only ranks, so it needs no score normalisation. A document that appears in four of five lists, even at rank 6, beats one that appears at rank 1 in a single list. That is the noise filter: each reformulation adds some junk, but junk rarely agrees across queries.
Relatives you should be able to name
- Multi-query retrieval — the same idea without RRF (results are just unioned). LangChain's
MultiQueryRetrieveris the classic example (since LangChain 1.0 it lives in thelangchain-classicpackage). - Query rewriting — one improved query instead of several; cheaper, less recall.
- HyDE — embed a generated answer instead of the question (covered in the hypothetical-questions lesson).
- Hybrid search — BM25 plus dense, fused with RRF. Often you run RAG Fusion over hybrid search.
When it does not help
- Precise lookups — an order ID, an error code, a clause number. The original query is already the best query, and extra ones add noise.
- The corpus lacks the answer. Five phrasings of a question about a missing policy still find nothing.
- Latency-critical paths — the generation call adds a few hundred milliseconds before any search starts.
A real-life example
An Indian telecom runs a customer-support bot over its help articles, which are written in English. Customers type in English, Hindi and Hinglish. A typical message:
recharge ho gaya par data nahi chal rahaSingle-query dense search returns articles about "recharge offers" and "payment failed" — the word recharge dominates. The article that answers it is titled "Data services not active after a successful top-up".
With RAG Fusion, a small model writes four queries:
1. Mobile data not working after successful recharge2. Data pack not activated after top-up3. Internet not working after prepaid plan payment4. How long does data activation take after recharge?Queries 2 and 4 both put the right article in their top three, so RRF lifts it to rank 1 overall. On a 500-message test set of real tickets, the team measures recall@10 before and after, per language. The gain is largest for Hinglish messages and close to zero for English queries that already match article titles. So they turn fusion on only when the language detector says Hindi or mixed, or when the query is under six words.
Follow-up questions to expect
- "How many reformulations?" — Start with 3–4. More adds cost and noise faster than recall; measure recall@k at each setting.
- "Why RRF and not average the similarity scores?" — Scores from different queries (and from BM25 versus dense) are not on the same scale. RRF needs only ranks, so it is robust without tuning.
- "Can the reformulations drift off-topic?" — Yes. Always keep the original query as one of the lists, and constrain the prompt to rephrasings, not new questions.