LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

Write a function to implement a hybrid search in LangChain.


Rank of the right chunk for two queriesU30 error code191money cut,order failed1212QueryBM25VectorHybrid
Each retriever fails where the other is strong, and rank fusion keeps the right chunk near the top for both.

What you need to know

Why combine them

QueryBM25Vectors
ERR_5012 timeoutFinds the exact error pageReturns generic timeout articles
"my money got deducted but order failed"Weak: no shared words with "payment reversal"Finds the reversal policy

Each method fails where the other is strong, so fusing them is more robust than either alone.

The function

Python
from langchain_community.retrievers import BM25Retriever      # pip install rank_bm25from langchain_classic.retrievers import EnsembleRetrieverdef hybrid_retriever(chunks, vector_store, k: int = 5, weights=(0.4, 0.6)):    keyword = BM25Retriever.from_documents(chunks, k=k)    dense = vector_store.as_retriever(search_kwargs={"k": k})    return EnsembleRetriever(retrievers=[keyword, dense],                             weights=list(weights), id_key="chunk_id")retriever = hybrid_retriever(chunks, vector_store)retriever.invoke("ERR_5012 timeout on checkout")
  • BM25Retriever.from_documents builds the keyword index in memory from the same chunks you embedded.
  • EnsembleRetriever runs both and applies reciprocal rank fusion: each document's score is the sum of weight / (60 + rank) over the lists it appears in (60 is the default constant c). Because only ranks are used, the BM25 score and cosine score never have to be compared.
  • id_key tells it which metadata field identifies the same chunk in both lists; otherwise it compares the text.
  • weights shift the balance. Code-heavy or id-heavy content favours BM25; conversational queries favour vectors.

Scaling beyond memory

BM25Retriever keeps every chunk in Python memory and must be rebuilt when chunks change. That is fine up to tens of thousands of chunks. Beyond that, use a backend with built-in hybrid search — Elasticsearch/OpenSearch, Qdrant or Weaviate (dense + sparse vectors), or Postgres with pgvector plus full-text tsvector — and fuse there.

A real-life example

A payments company's developer-docs assistant fails on queries like "What does error U30 mean?" — a UPI error code. Vector search returns general "payment failed" pages, because to an embedding model U30 is noise.

With hybrid search at weights 0.5/0.5, the exact error-code page ranks first because BM25 matches U30 precisely, and vector results still cover questions phrased in plain words. On their 200-question test set, recall@5 goes from 0.71 (vectors only) and 0.66 (BM25 only) to 0.84 hybrid. Six months later the docs grow to 400,000 chunks, and they move to OpenSearch's hybrid query so the keyword index no longer lives in each Python worker.

Follow-up questions to expect

  • "Why reciprocal rank fusion rather than adding scores?" — BM25 scores are unbounded and cosine scores are 0–1; adding them lets one method dominate by accident.
  • "How do you set the weights?" — Grid-search a few pairs (0.3/0.7, 0.5/0.5, 0.7/0.3) on a labelled set.
  • "What is sparse-dense search?" — The same idea inside one store: a sparse vector (keyword-like) and a dense embedding per chunk, scored together.