Course Content
LangChain Mastery
7 sections · 109 lessons
How do you implement a LangChain retriever with metadata filtering?
What you need to know
Step 1: metadata at index time
1from langchain_core.documents import Document23docs = [4 Document(page_content=text,5 metadata={"tenant_id": "acme", "dept": "hr", "year": 2026, "lang": "en"})6 for text in texts7]8vector_store.add_documents(docs)Plan the fields before indexing. Adding a field later means re-writing every chunk.
Step 2: filter at query time
1retriever = vector_store.as_retriever(search_kwargs={2 "k": 4,3 "filter": {"$and": [{"dept": "hr"}, {"year": {"$gte": 2025}}]}, # Chroma syntax4})The store applies the filter inside the search, so only matching vectors compete for the top k.
Syntax is not portable
| Store | Filter looks like |
|---|---|
| Chroma | {"$and": [{"dept": "hr"}, {"year": {"$gte": 2025}}]} |
InMemoryVectorStore | a Python function: filter=lambda doc: doc.metadata["dept"] == "hr" |
pgvector (langchain-postgres) | dict with operators such as $eq, $in over a JSONB column |
| Pinecone, Qdrant | their own dict or filter-object formats |
Always check the integration's docs; a wrong key often returns zero results rather than an error.
Per-request filters
Create the retriever inside the request, from trusted data:
def retriever_for(user): return vector_store.as_retriever(search_kwargs={ "k": 4, "filter": {"tenant_id": user.tenant_id}})In an agent, a retrieval tool can read the tenant from runtime context (a ToolRuntime parameter, filled from context= on agent.invoke), so the model never sees or controls the value.
Pre-filter vs post-filter
Some stores filter before the nearest-neighbour search (correct k results); others search first and filter after, which can return fewer than k results — or none — when the filter is selective. Know which one your store does.
A real-life example
A SaaS company runs one HR assistant for 300 client companies in a single vector index. In a test, an employee of client A asked "What is the notice period?" and got client B's policy — the only filter was a sentence in the prompt: "Only use documents from the user's company."
The fix: every chunk gets tenant_id metadata, and the retriever is created per request with filter={"tenant_id": session.tenant_id} from the login token. A regression test now asks the same question as users of two tenants and asserts that no chunk from the other tenant is ever returned. Search is also faster, because each query scans about 1/300th of the index.
Follow-up questions to expect
- "Why not trust the prompt to restrict sources?" — The model already saw the forbidden text once it is in the prompt; instructions reduce leaks but cannot guarantee zero.
- "When would you use separate indexes per tenant instead?" — For strict isolation or very large tenants; filtering one shared index is simpler for many small tenants.
- "How do filters interact with k?" — With post-filtering you may get fewer than k results; raise
fetch_kor use a pre-filtering store.