Course Content
Advanced RAG
3 sections · 38 lessons
How does Reciprocal Rank Fusion (RRF) combine retrieval results?
What you need to know
Why scores cannot simply be added
- A BM25 score is unbounded and depends on the corpus and the query length: 7.2 on one query, 31.5 on another.
- A cosine similarity sits between -1 and 1, and most results cluster in a narrow band such as 0.70–0.85.
Adding them lets BM25 dominate. Normalising each list to 0–1 depends on each list's spread and changes after re-indexing. RRF sidesteps all of it: rank 1 is rank 1 in any list.
The formula, in code
1from collections import defaultdict23def rrf(rankings, k=60, weights=None):4 weights = weights or [1.0] * len(rankings)5 scores = defaultdict(float)6 for ranking, w in zip(rankings, weights):7 for rank, doc_id in enumerate(ranking, start=1):8 scores[doc_id] += w / (k + rank)9 return sorted(scores.items(), key=lambda x: x[1], reverse=True)1011bm25 = ["KYC-114", "KYC-007", "AML-220", "FAQ-031"]12dense = ["AML-220", "KYC-090", "KYC-114", "FAQ-031"]13for doc, s in rrf([bm25, dense]):14 print(f"{doc:8} {s:.4f}")KYC-114 0.0323AML-220 0.0323FAQ-031 0.0312KYC-007 0.0161KYC-090 0.0161Read the output carefully. KYC-114 (ranks 1 and 3) and AML-220 (ranks 3 and 1) tie at the top. FAQ-031 is only 4th in both lists, yet it beats KYC-007, which was 2nd in BM25 but missing from the dense list. Agreement across lists beats a single high rank — that is RRF's built-in noise filter.
What k does
- Small
k(e.g. 1–10): the gap between rank 1 and rank 2 is large, so each list's top hits dominate. - Large
k(60, the value from the original 2009 paper): the curve is flat, so appearing in many lists matters more than topping one.
In practice, 60 works well and is rarely worth tuning. Per-list weights are more useful: if BM25 is clearly stronger for your ID-heavy queries, give it 1.5.
Where you meet RRF in 2026
- Most search engines and vector databases offer hybrid search with RRF built in (Elasticsearch, OpenSearch, Qdrant, Weaviate, and others), so you rarely write it yourself.
- RAG Fusion and federated RAG use it to merge lists from several queries or sources.
- It is a first-stage fusion. A cross-encoder reranker usually follows it and produces the one real relevance score.
A real-life example
A bank's compliance assistant has thousands of documents with codes like KYC-114 and AML-220. Officers mix two kinds of query:
- "What does KYC-114 say about re-verification?" — BM25 finds
KYC-114at rank 1 immediately; the dense retriever puts it at rank 9, because the code carries little meaning for the embedding. - "What do we do when a customer refuses to update their address?" — dense retrieval finds the right policy; BM25 returns documents that happen to contain "update" and "address" many times.
Dense-only search fails the first type; BM25-only fails the second. With both retrievers returning a top 50, fused by RRF and followed by a reranker, both types succeed. On a 400-question test set, the team measures recall@10 for dense-only, BM25-only and hybrid, split by "contains a document code" versus "natural language". Hybrid is best or tied-best in both groups, which neither single retriever manages. They then try a weight of 1.3 on BM25 and keep it, because it helps code-heavy queries without hurting the others.
Follow-up questions to expect
- "When would you use score-based fusion instead?" — When both retrievers produce calibrated scores and you have labelled data to tune a weighted combination. Some engines offer normalised-score fusion; it can beat RRF when tuned, but it is more fragile.
- "Does RRF improve precision at the top?" — Mainly recall and robustness. Precision at the top comes from the reranker that follows.
- "How many results from each list?" — Enough that the right document is usually present in at least one list, commonly 20–100 each; measure recall@N to pick it.