Course Content
LangChain Mastery
7 sections · 109 lessons
Implement a function to combine multiple vector stores in LangChain.
What you need to know
Why you end up with several stores
- Different teams own different content (docs, FAQs, tickets).
- Different embedding models suit different content (code vs prose).
- Different access rules or retention (public docs vs internal notes).
You cannot merge vectors from different embedding models into one index, so combining at retrieval time is the normal answer.
Option 1: fuse rankings with EnsembleRetriever
1from langchain_classic.retrievers import EnsembleRetriever23combined = EnsembleRetriever(4 retrievers=[docs_store.as_retriever(search_kwargs={"k": 5}),5 faq_store.as_retriever(search_kwargs={"k": 5})],6 weights=[0.6, 0.4],7 id_key="id", # metadata field used to spot the same doc in both lists8)9results = combined.invoke("how to reset UPI PIN")In LangChain 1.x, EnsembleRetriever is imported from langchain_classic.retrievers.
Option 2: run in parallel, keep every source
1from langchain_core.runnables import RunnableParallel, RunnableLambda23def combine_stores(stores: dict, k: int = 4):4 """Query several vector stores in parallel, tag each hit, drop duplicates."""5 branches = {name: s.as_retriever(search_kwargs={"k": k}) for name, s in stores.items()}67 def merge(results: dict) -> list:8 seen, merged = set(), []9 for name, docs in results.items():10 for d in docs:11 key = d.metadata.get("id") or d.page_content12 if key not in seen:13 seen.add(key)14 d.metadata["store"] = name15 merged.append(d)16 return merged1718 return RunnableParallel(branches) | RunnableLambda(merge)1920retriever = combine_stores({"docs": docs_store, "tickets": ticket_store})RunnableParallel runs the retrievers concurrently, so latency is about the slowest store, not the sum. The store tag lets the prompt and the UI say where each fact came from. MergerRetriever (also in langchain_classic.retrievers) does a similar interleaved concatenation.
What to watch
- Context size — 3 stores × k=5 is 15 chunks. Cap the merged list or rerank it down to 5–6.
- Duplicates — the same article indexed in two stores. Dedupe by a stable id, or with
EmbeddingsRedundantFilterfromlangchain_community.document_transformers. - Latency — one slow store delays everything; give each retriever a timeout or a fallback to an empty list.
A real-life example
A bank's customer-support assistant has three stores: public help articles, internal agent notes, and resolved tickets. First version: one EnsembleRetriever over all three. Complaints followed — for "how do I reset my UPI PIN", the tickets store dominated with long, messy chats and the clean help article dropped to rank 6.
The team switched to the parallel approach: 3 help articles, 2 agent notes and 1 ticket, each tagged. The prompt says "prefer help articles; use tickets only as examples". The average prompt shrank from 15 chunks to 6, cost per answer fell by about half, and the help article was always present.
Follow-up questions to expect
- "How do you choose the weights?" — Start equal, then tune on a labelled question set; weights are a guess until measured.
- "Could you just put everything in one store?" — Yes, with a
sourcemetadata field and filters, if all content uses the same embedding model and access rules. That is often simpler. - "Why not average the raw scores?" — Different stores and models produce scores on different scales; rank fusion avoids comparing them.