Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Scenario – 6: Hybrid Retrieval Architecture
Scenario: you need semantic search plus exact keyword matching in a LangChain retrieval stack, because pure vector search misses error codes and product SKUs. How do you design it?
What you need to know
Dense (embedding) search
- Matches meaning: "can't log in" finds "authentication failure"
- Blurs exact tokens: codes, SKUs, version numbers
- Works across paraphrases and some languages
Sparse (BM25 keyword) search
- Matches exact words and rare tokens precisely
- Misses synonyms and paraphrases
- Cheap, explainable and fast
Each misses what the other catches, which is why combining them is a strong default for technical corpora.
The composition in LangChain
1from langchain_classic.retrievers import EnsembleRetriever, ContextualCompressionRetriever2from langchain_classic.retrievers.document_compressors import CrossEncoderReranker3from langchain_community.retrievers import BM25Retriever45hybrid = EnsembleRetriever(6 retrievers=[vectorstore.as_retriever(search_kwargs={"k": 30}),7 BM25Retriever.from_documents(docs, k=30)],8 weights=[0.5, 0.5])9retriever = ContextualCompressionRetriever(10 base_compressor=CrossEncoderReranker(model=cross_encoder, top_n=5),11 base_retriever=hybrid)EnsembleRetriever fuses the two ranked lists by rank, not by score. That is the right choice because a BM25 score of 14.2 and a cosine of 0.81 are on different scales; rank-based fusion avoids normalising them.
Scaling beyond memory
BM25Retriever builds an in-memory index from the documents each time the process starts. That is fine up to tens of thousands of documents, but it doesn't scale or persist, and it must be rebuilt when documents change. At real scale, use native hybrid search in the store:
| Store | Hybrid approach |
|---|---|
| Qdrant | Dense plus sparse vectors in one collection, fused server-side |
| Weaviate | Built-in hybrid query with a weighting parameter |
| Elasticsearch or OpenSearch | BM25 plus kNN in one query |
| Postgres | tsvector full-text search alongside pgvector, fused in SQL |
One query then does both legs, and nothing is rebuilt at startup.
Tuning
- Label — 200 queries with their correct documents, including code and SKU lookups.
- Start at 50/50 — measure recall@30 for the ensemble.
- Sweep weights — keyword-heavy corpora often do better leaning towards the sparse side.
- Rerank — measure precision@5 after the cross-encoder, and its latency per stage.
A real-life example
Scenario (illustrative numbers). An industrial-equipment maker's service assistant helps technicians with manuals and fault codes. With dense-only search, queries like "E-217 on the CX-400 compressor" return pages about other fault codes on similar models; recall@30 on 250 labelled queries is 0.64, and just 0.38 for code-style queries.
Adding BM25 through EnsembleRetriever lifts overall recall@30 to 0.89 and code-style recall to 0.93. The best weights on their data are 0.4 dense and 0.6 sparse. With a reranker, precision@5 reaches 0.82. As the manual library grows past 80,000 documents, they move both legs into their vector database's native hybrid query and drop the in-memory BM25 index.
Follow-up questions to expect
- "Why reciprocal rank fusion instead of adding scores?" — Scores from BM25 and cosine similarity are not comparable; fusing by rank is simple and robust without calibration.
- "What about learned sparse models like SPLADE?" — They add term expansion to keyword search and can beat plain BM25; some vector databases store them as sparse vectors.
- "Does hybrid search help with multilingual queries?" — The BM25 leg doesn't cross languages; translate the query for that leg, or rely on a multilingual dense model.