LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

How do you implement a LangChain retriever with metadata filtering?


What you need to know

Step 1: metadata at index time

Python
from langchain_core.documents import Documentdocs = [    Document(page_content=text,             metadata={"tenant_id": "acme", "dept": "hr", "year": 2026, "lang": "en"})    for text in texts]vector_store.add_documents(docs)

Plan the fields before indexing. Adding a field later means re-writing every chunk.

Step 2: filter at query time

Python
retriever = vector_store.as_retriever(search_kwargs={    "k": 4,    "filter": {"$and": [{"dept": "hr"}, {"year": {"$gte": 2025}}]},   # Chroma syntax})

The store applies the filter inside the search, so only matching vectors compete for the top k.

Syntax is not portable

StoreFilter looks like
Chroma{"$and": [{"dept": "hr"}, {"year": {"$gte": 2025}}]}
InMemoryVectorStorea Python function: filter=lambda doc: doc.metadata["dept"] == "hr"
pgvector (langchain-postgres)dict with operators such as $eq, $in over a JSONB column
Pinecone, Qdranttheir own dict or filter-object formats

Always check the integration's docs; a wrong key often returns zero results rather than an error.

Per-request filters

Create the retriever inside the request, from trusted data:

Python
def retriever_for(user):    return vector_store.as_retriever(search_kwargs={        "k": 4, "filter": {"tenant_id": user.tenant_id}})

In an agent, a retrieval tool can read the tenant from runtime context (a ToolRuntime parameter, filled from context= on agent.invoke), so the model never sees or controls the value.

Pre-filter vs post-filter

Some stores filter before the nearest-neighbour search (correct k results); others search first and filter after, which can return fewer than k results — or none — when the filter is selective. Know which one your store does.

A real-life example

A SaaS company runs one HR assistant for 300 client companies in a single vector index. In a test, an employee of client A asked "What is the notice period?" and got client B's policy — the only filter was a sentence in the prompt: "Only use documents from the user's company."

The fix: every chunk gets tenant_id metadata, and the retriever is created per request with filter={"tenant_id": session.tenant_id} from the login token. A regression test now asks the same question as users of two tenants and asserts that no chunk from the other tenant is ever returned. Search is also faster, because each query scans about 1/300th of the index.

Follow-up questions to expect

  • "Why not trust the prompt to restrict sources?" — The model already saw the forbidden text once it is in the prompt; instructions reduce leaks but cannot guarantee zero.
  • "When would you use separate indexes per tenant instead?" — For strict isolation or very large tenants; filtering one shared index is simpler for many small tenants.
  • "How do filters interact with k?" — With post-filtering you may get fewer than k results; raise fetch_k or use a pre-filtering store.