Course Content
RAG Systems
12 sections · 66 lessons
How do metadata filters improve retrieval quality?
What you need to know
The superseded-policy problem, in a real run
An HR store holds the current handbook (v7: "bands 1 to 3 serve 60 days; band 4 and above serve 90 days") and an old one (v5: "all employees must serve 30 days"). The query "What notice period do I have to serve?" returned:
0.158 hr/handbook-v5.md Notice period <- old, wrong, ranked first0.189 hr/handbook-v7.md Notice periodThe old version is shorter and simpler, so it is closer in meaning to the question. No embedding model or prompt can know it was superseded. A filter can:
1store.similarity_search(2 "What notice period do I have to serve?", k=2,3 filter={"$and": [{"country": "IN"}, {"year": {"$gte": 2025}}]},4)5# -> hr/handbook-v7.md Notice period, hr/handbook-v7.md Notice buy-outSyntax traps, from real runs on Chroma
- Multiple conditions need
$and.filter={"country": "IN", "year": 2026}raisesValueError: Expected where to have exactly one operator. Other stores accept a plain dict; check yours. - Values are case-sensitive.
filter={"country": "in"}returned an empty list with no error, because the data says"IN". Normalise values at ingest (for example, lowercase everything) and at query time.
Pre-filtering and post-filtering
Pre-filter (during search)
- Filter applied while walking the index
- Always returns k results if k exist
- Supported by Qdrant, pgvector, Weaviate, Milvus and others
- Needs filter-aware index traversal
Post-filter (after search)
- Search top k first, then drop non-matching
- Can return fewer than k, or none
- Common in simple FAISS wrappers
- Fix: fetch more, then filter
If a strict filter matches 1% of the corpus and you post-filter the top 10, you will often get zero results.
Security is a filter, not a prompt
For access control, store the allowed groups on every chunk at ingest (for example acl: ["hr", "managers"]) and filter on the user's groups at query time. Never retrieve everything and tell the model "do not reveal manager-only content"; a document already in the prompt can leak through a summary, a quote or a prompt-injection attack.
Filters from the question
Some filters come from the user, not the session: "What did the 2024 Vendor A contract say?". A self-querying step asks an LLM to turn that into a structured filter ({"vendor": "vendor-a", "year": 2024}) plus a search query. Validate the generated filter against allowed fields before running it.
A real-life example
A legal-contract search tool serves several business units. Each contract chunk carries business_unit, counterparty, effective_date, status (active, expired, draft) and acl.
An associate in the retail unit asks, "What is our termination notice with the logistics vendor?". The retriever applies status = active, acl contains the user's groups, and counterparty extracted from the question. Before these filters, the top result was an expired 2021 contract with the same vendor, and sometimes a draft that was never signed. After them, the answer came from the active 2024 contract only, and the tool could no longer show wholesale-unit contracts to retail users.
Follow-up questions to expect
- "What if the filter removes everything?" — Tell the user nothing matched their scope, or relax non-security filters (such as year) and say so. Never relax an access-control filter.
- "How do you handle document versions?" — Store
versionandeffective_from/effective_to, and filter to the current one by default; allow explicit "as of" questions to override. - "Do filters slow down search?" — Selective filters can speed it up; very selective ones can hurt HNSW recall, and good engines then switch to exact search on the filtered set.