Advanced RAG

Course Content

Advanced RAG

3 sections · 38 lessons

How does Reciprocal Rank Fusion (RRF) combine retrieval results?


Ranks in, one fused order out (k = 60)130.0323310.0323440.03122absent0.0161absent20.0161BM25 rankDense rankRRF scoreKYC-114AML-220FAQ-031KYC-007KYC-090
Fourth in both lists beats second in only one — agreement across retrievers is RRF's built-in noise filter.

What you need to know

Why scores cannot simply be added

  • A BM25 score is unbounded and depends on the corpus and the query length: 7.2 on one query, 31.5 on another.
  • A cosine similarity sits between -1 and 1, and most results cluster in a narrow band such as 0.70–0.85.

Adding them lets BM25 dominate. Normalising each list to 0–1 depends on each list's spread and changes after re-indexing. RRF sidesteps all of it: rank 1 is rank 1 in any list.

The formula, in code

Python
from collections import defaultdictdef rrf(rankings, k=60, weights=None):    weights = weights or [1.0] * len(rankings)    scores = defaultdict(float)    for ranking, w in zip(rankings, weights):        for rank, doc_id in enumerate(ranking, start=1):            scores[doc_id] += w / (k + rank)    return sorted(scores.items(), key=lambda x: x[1], reverse=True)bm25  = ["KYC-114", "KYC-007", "AML-220", "FAQ-031"]dense = ["AML-220", "KYC-090", "KYC-114", "FAQ-031"]for doc, s in rrf([bm25, dense]):    print(f"{doc:8} {s:.4f}")
Text
KYC-114  0.0323AML-220  0.0323FAQ-031  0.0312KYC-007  0.0161KYC-090  0.0161

Read the output carefully. KYC-114 (ranks 1 and 3) and AML-220 (ranks 3 and 1) tie at the top. FAQ-031 is only 4th in both lists, yet it beats KYC-007, which was 2nd in BM25 but missing from the dense list. Agreement across lists beats a single high rank — that is RRF's built-in noise filter.

What k does

  • Small k (e.g. 1–10): the gap between rank 1 and rank 2 is large, so each list's top hits dominate.
  • Large k (60, the value from the original 2009 paper): the curve is flat, so appearing in many lists matters more than topping one.

In practice, 60 works well and is rarely worth tuning. Per-list weights are more useful: if BM25 is clearly stronger for your ID-heavy queries, give it 1.5.

Where you meet RRF in 2026

  • Most search engines and vector databases offer hybrid search with RRF built in (Elasticsearch, OpenSearch, Qdrant, Weaviate, and others), so you rarely write it yourself.
  • RAG Fusion and federated RAG use it to merge lists from several queries or sources.
  • It is a first-stage fusion. A cross-encoder reranker usually follows it and produces the one real relevance score.

A real-life example

A bank's compliance assistant has thousands of documents with codes like KYC-114 and AML-220. Officers mix two kinds of query:

  • "What does KYC-114 say about re-verification?" — BM25 finds KYC-114 at rank 1 immediately; the dense retriever puts it at rank 9, because the code carries little meaning for the embedding.
  • "What do we do when a customer refuses to update their address?" — dense retrieval finds the right policy; BM25 returns documents that happen to contain "update" and "address" many times.

Dense-only search fails the first type; BM25-only fails the second. With both retrievers returning a top 50, fused by RRF and followed by a reranker, both types succeed. On a 400-question test set, the team measures recall@10 for dense-only, BM25-only and hybrid, split by "contains a document code" versus "natural language". Hybrid is best or tied-best in both groups, which neither single retriever manages. They then try a weight of 1.3 on BM25 and keep it, because it helps code-heavy queries without hurting the others.

Follow-up questions to expect

  • "When would you use score-based fusion instead?" — When both retrievers produce calibrated scores and you have labelled data to tune a weighted combination. Some engines offer normalised-score fusion; it can beat RRF when tuned, but it is more fragile.
  • "Does RRF improve precision at the top?" — Mainly recall and robustness. Precision at the top comes from the reranker that follows.
  • "How many results from each list?" — Enough that the right document is usually present in at least one list, commonly 20–100 each; measure recall@N to pick it.