Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your vector DB stores menus from 5,000 restaurants. Users need semantic search AND hard filters (city, price range). Pure vector search doesn't support filters. How do you design hybrid search?
What you need to know
Pre-filter, post-filter, filtered search
Post-filter
- Search all vectors, then drop non-matches
- Top 10 overall may contain zero from Nagpur
- Returns too few results on selective filters
- Simple, but wrong for this use case
Filtered search
- Filter is applied while walking the index
- Only matching items can be returned
- Needs indexed payload fields
- The right default for hard filters
Engines differ in how they do it. Qdrant applies payload filters during HNSW traversal and switches to an exact scan over the filtered set when very few points match. pgvector (0.8 and later) can keep scanning the HNSW index until enough rows pass the WHERE clause (iterative scans). Know your engine's behaviour, and set its thresholds deliberately.
Model the data for filtering
Each menu-item chunk carries its restaurant's fields, because filters apply to chunks, not to parent records:
1point = {2 "id": "rest_381:item_12",3 "vector": embed("Kadai paneer — cottage cheese in a spicy tomato and capsicum gravy"),4 "payload": {"restaurant_id": 381, "city": "pune", "price": 320, # numeric, for ranges5 "cuisine": ["north indian"], "veg": True, "is_open": True, "rating": 4.3},6}78hits = client.query_points("menus", query=qvec, limit=50, query_filter=Filter(must=[9 FieldCondition(key="city", match=MatchValue(value="pune")),10 FieldCondition(key="price", range=Range(lte=400)),11 FieldCondition(key="is_open", match=MatchValue(value=True)),12])).pointsCreate payload indexes on the filtered fields; without them, filtering is slow. Keep price as a number so range queries work.
The lexical half, and fusion
Embeddings are weak on rare proper nouns — dish names ("kadai paneer"), restaurant names, brands. Run BM25 with the same filters, then merge the lists with Reciprocal Rank Fusion, which uses rank positions, so the two incompatible score scales never need to be compared.
RRF score(item) = sum over lists of 1 / (k + rank of item in that list), with k = 60Finally, rerank the fused top 50 with a cross-encoder and return 10.
Watch the selective queries
| Metric | Why |
|---|---|
| Recall@10 on a filtered golden set | Quality where filters apply |
| Filter selectivity distribution | How often queries match very few items |
| p95 latency by selectivity bucket | The slowest queries are usually the most selective |
A real-life example
Scenario, numbers made up. A food-ordering app searches 1.2M menu items from 5,000 restaurants. The first version runs vector search for the top 50, then filters by city and price. Users in smaller cities often get "no results" for "cheap veg biryani", because the global top 50 is full of Bengaluru and Mumbai items.
The team moves to filtered search with payload indexes on city, price, veg and open status, adds BM25 on dish names, fuses with RRF and reranks. Zero-result searches in tier-2 cities drop from 18% to under 2%. A search for "Kadai Paneer" now puts exact dish matches first instead of generic "paneer curry" items, and p95 latency stays under 120 ms.
Follow-up questions to expect
- "Why not just normalise the two scores and add them?" — BM25 and cosine scores have different, query-dependent scales, so a fixed weighting behaves differently for every query. RRF avoids that by using ranks.
- "How do you handle 'near me'?" — Store a geo point and use the engine's geo-radius filter, or pre-compute a delivery-zone ID.
- "What if a filter matches only 20 items?" — Then an exact scan of those 20 is fastest and perfectly accurate; good engines do this automatically below a threshold.