Course Content
RAG Systems
12 sections · 66 lessons
What is hybrid search and why is it useful in RAG systems?
What you need to know
BM25 in one paragraph
BM25 needs no model and no GPU. Its weakness is vocabulary: "send back headphones" shares no words with "returns: earbuds can be returned".
Reciprocal rank fusion (RRF)
RRF(d) = sum over each result list of 1 / (k + rank of d in that list), k = 60A document found by both methods gets two contributions; a document found by only one gets one. The constant 60 comes from the original 2009 paper by Cormack and colleagues and keeps any single top rank from dominating.
A real run, fused by hand
A bank's card-help corpus has 8 short passages. For the query "my card stopped working after I typed the wrong code a few times", a real run of rank_bm25 and bge-small-en-v1.5 gave these top-5 lists (passage IDs):
BM25: 0 (PIN blocked), 4 (refunds), 3 (merchant declined), 5 (lost card), 2 (daily limit)Dense: 0 (PIN blocked), 3, 5, 1 (card expired), 2BM25 ranked the refunds passage second only because it contains "working" ("5 to 7 working days"). Fusing:
| Passage | BM25 rank | Dense rank | RRF score |
|---|---|---|---|
| 0 PIN blocked | 1 | 1 | 1/61 + 1/61 = 0.0328 |
| 3 Merchant declined | 3 | 2 | 1/63 + 1/62 = 0.0320 |
| 5 Lost card | 4 | 3 | 1/64 + 1/63 = 0.0315 |
| 2 Daily limit | 5 | 5 | 1/65 + 1/65 = 0.0308 |
| 4 Refunds | 2 | — | 1/62 = 0.0161 |
The false keyword match drops from 2nd to 5th, because only one method liked it. For "what does E-4012 mean", BM25 gave the E-4012 passage a score of 1.54 and every other passage 0, a decisive exact match; dense search also ranked it first here, but by a narrower margin (0.666 against 0.521). On a larger catalogue with thousands of similar codes, that margin is where dense search starts to slip.
The code
1import re2import numpy as np3from rank_bm25 import BM25Okapi45tokenize = lambda s: re.findall(r"[a-z0-9-]+", s.lower()) # keeps "e-4012" whole6bm25 = BM25Okapi([tokenize(d) for d in docs])78def rrf(rankings: list[list[int]], k: int = 60) -> list[int]:9 scores: dict[int, float] = {}10 for ranking in rankings:11 for rank, doc_id in enumerate(ranking, start=1):12 scores[doc_id] = scores.get(doc_id, 0) + 1 / (k + rank)13 return sorted(scores, key=scores.get, reverse=True)1415sparse = np.argsort(-bm25.get_scores(tokenize(query)))[:20].tolist()16dense = np.argsort(-(D @ q_vec))[:20].tolist() # D, q_vec: normalised embeddings17fused = rrf([sparse, dense])The tokenizer matters: splitting "E-4012" into "e" and "4012" would destroy the exact match that makes BM25 useful.
In practice
- LangChain:
EnsembleRetriever(inlangchain_classic.retrievers) over aBM25Retrieverand a vector retriever runs weighted RRF withc = 60. - Native engines: Elasticsearch and OpenSearch, Qdrant, Weaviate, Milvus, Vespa and Postgres (
pgvectorplus full-text search) all support some form of hybrid query. - Learned sparse models (SPLADE, or the sparse output of BGE-M3) replace BM25 with learned term weights that also expand synonyms.
- Contextual BM25: index the chunk with its prepended context line, so keyword search also benefits from contextual retrieval.
A real-life example
A bank's FAQ bot gets many questions that quote error codes from the app ("what is E-4021?"). With dense search alone, answers sometimes described a neighbouring code, because the embedding saw "E-4021" and "E-4012" as nearly the same. Customers followed the wrong fix.
Adding BM25 with a tokenizer that keeps codes whole, fused with RRF, made the exact code's passage appear first for these queries, while paraphrased questions ("my card got locked after wrong PINs") still worked through the dense side. The team kept 30 fused candidates for a reranker and sent the top 4.
Follow-up questions to expect
- "RRF or a weighted score blend?" — RRF needs no score normalisation and is robust. A weighted blend (
alpha × dense + (1 − alpha) × sparse) can be tuned further, but only after normalising both scores to the same range. - "How do you keep the two indexes in sync?" — Write both in the same ingestion job with the same chunk IDs, or use one engine that stores both.
- "Does BM25 work for Hindi?" — Yes, with a tokenizer suited to the script; stemming and transliteration need extra care.