Course Content
LangChain Mastery
7 sections · 109 lessons
Write a function to implement a hybrid search in LangChain.
What you need to know
Why combine them
| Query | BM25 | Vectors |
|---|---|---|
ERR_5012 timeout | Finds the exact error page | Returns generic timeout articles |
| "my money got deducted but order failed" | Weak: no shared words with "payment reversal" | Finds the reversal policy |
Each method fails where the other is strong, so fusing them is more robust than either alone.
The function
1from langchain_community.retrievers import BM25Retriever # pip install rank_bm252from langchain_classic.retrievers import EnsembleRetriever34def hybrid_retriever(chunks, vector_store, k: int = 5, weights=(0.4, 0.6)):5 keyword = BM25Retriever.from_documents(chunks, k=k)6 dense = vector_store.as_retriever(search_kwargs={"k": k})7 return EnsembleRetriever(retrievers=[keyword, dense],8 weights=list(weights), id_key="chunk_id")910retriever = hybrid_retriever(chunks, vector_store)11retriever.invoke("ERR_5012 timeout on checkout")BM25Retriever.from_documentsbuilds the keyword index in memory from the same chunks you embedded.EnsembleRetrieverruns both and applies reciprocal rank fusion: each document's score is the sum ofweight / (60 + rank)over the lists it appears in (60 is the default constantc). Because only ranks are used, the BM25 score and cosine score never have to be compared.id_keytells it which metadata field identifies the same chunk in both lists; otherwise it compares the text.weightsshift the balance. Code-heavy or id-heavy content favours BM25; conversational queries favour vectors.
Scaling beyond memory
BM25Retriever keeps every chunk in Python memory and must be rebuilt when chunks change. That is fine up to tens of thousands of chunks. Beyond that, use a backend with built-in hybrid search — Elasticsearch/OpenSearch, Qdrant or Weaviate (dense + sparse vectors), or Postgres with pgvector plus full-text tsvector — and fuse there.
A real-life example
A payments company's developer-docs assistant fails on queries like "What does error U30 mean?" — a UPI error code. Vector search returns general "payment failed" pages, because to an embedding model U30 is noise.
With hybrid search at weights 0.5/0.5, the exact error-code page ranks first because BM25 matches U30 precisely, and vector results still cover questions phrased in plain words. On their 200-question test set, recall@5 goes from 0.71 (vectors only) and 0.66 (BM25 only) to 0.84 hybrid. Six months later the docs grow to 400,000 chunks, and they move to OpenSearch's hybrid query so the keyword index no longer lives in each Python worker.
Follow-up questions to expect
- "Why reciprocal rank fusion rather than adding scores?" — BM25 scores are unbounded and cosine scores are 0–1; adding them lets one method dominate by accident.
- "How do you set the weights?" — Grid-search a few pairs (0.3/0.7, 0.5/0.5, 0.7/0.3) on a labelled set.
- "What is sparse-dense search?" — The same idea inside one store: a sparse vector (keyword-like) and a dense embedding per chunk, scored together.