Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your vector DB stores menus from 5,000 restaurants. Users need semantic search AND hard filters (city, price range). Pure vector search doesn't support filters. How do you design hybrid search?


Filters inside both searches, then fusionFiltered vectorsearch: city, priceFiltered BM25for dish namesReciprocal RankFusion, k of 60Cross-encoderrerank,return 10Restaurant fields are copied onto every menu-item chunk.
Filtering during the search, not after it, is what took zero-result searches in smaller cities from 18 percent to under 2.

What you need to know

Pre-filter, post-filter, filtered search

Post-filter

  • Search all vectors, then drop non-matches
  • Top 10 overall may contain zero from Nagpur
  • Returns too few results on selective filters
  • Simple, but wrong for this use case

Filtered search

  • Filter is applied while walking the index
  • Only matching items can be returned
  • Needs indexed payload fields
  • The right default for hard filters

Engines differ in how they do it. Qdrant applies payload filters during HNSW traversal and switches to an exact scan over the filtered set when very few points match. pgvector (0.8 and later) can keep scanning the HNSW index until enough rows pass the WHERE clause (iterative scans). Know your engine's behaviour, and set its thresholds deliberately.

Model the data for filtering

Each menu-item chunk carries its restaurant's fields, because filters apply to chunks, not to parent records:

Python
point = {    "id": "rest_381:item_12",    "vector": embed("Kadai paneer — cottage cheese in a spicy tomato and capsicum gravy"),    "payload": {"restaurant_id": 381, "city": "pune", "price": 320,      # numeric, for ranges                "cuisine": ["north indian"], "veg": True, "is_open": True, "rating": 4.3},}hits = client.query_points("menus", query=qvec, limit=50, query_filter=Filter(must=[    FieldCondition(key="city", match=MatchValue(value="pune")),    FieldCondition(key="price", range=Range(lte=400)),    FieldCondition(key="is_open", match=MatchValue(value=True)),])).points

Create payload indexes on the filtered fields; without them, filtering is slow. Keep price as a number so range queries work.

The lexical half, and fusion

Embeddings are weak on rare proper nouns — dish names ("kadai paneer"), restaurant names, brands. Run BM25 with the same filters, then merge the lists with Reciprocal Rank Fusion, which uses rank positions, so the two incompatible score scales never need to be compared.

Text
RRF score(item) = sum over lists of 1 / (k + rank of item in that list),  with k = 60

Finally, rerank the fused top 50 with a cross-encoder and return 10.

Watch the selective queries

MetricWhy
Recall@10 on a filtered golden setQuality where filters apply
Filter selectivity distributionHow often queries match very few items
p95 latency by selectivity bucketThe slowest queries are usually the most selective

A real-life example

Scenario, numbers made up. A food-ordering app searches 1.2M menu items from 5,000 restaurants. The first version runs vector search for the top 50, then filters by city and price. Users in smaller cities often get "no results" for "cheap veg biryani", because the global top 50 is full of Bengaluru and Mumbai items.

The team moves to filtered search with payload indexes on city, price, veg and open status, adds BM25 on dish names, fuses with RRF and reranks. Zero-result searches in tier-2 cities drop from 18% to under 2%. A search for "Kadai Paneer" now puts exact dish matches first instead of generic "paneer curry" items, and p95 latency stays under 120 ms.

Follow-up questions to expect

  • "Why not just normalise the two scores and add them?" — BM25 and cosine scores have different, query-dependent scales, so a fixed weighting behaves differently for every query. RRF avoids that by using ranks.
  • "How do you handle 'near me'?" — Store a geo point and use the engine's geo-radius filter, or pre-compute a delivery-zone ID.
  • "What if a filter matches only 20 items?" — Then an exact scan of those 20 is fastest and perfectly accurate; good engines do this automatically below a threshold.