Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Scenario – 2: Missing Context Retrieval


What you need to know

The scenario: the answer exists in the documents, but the system says it cannot find it or answers from the wrong context.

Diagnose by layer

  1. Build the golden set — 100–200 real queries, each with its correct chunk labelled.
  2. Measure recall at 5, 20 and 100 — where does the correct chunk land?
  3. Classify each failure — missing from top 100, ranked low, or retrieved but ignored.
  4. Fix the biggest bucket — then re-measure.
Where the gold chunk isBroken layerTypical fixes
Not in top 100RetrievalQuery rewriting, hybrid BM25, chunking, filters
In top 100, below top 5RankingCross-encoder reranker over top 50
In top 5, answer ignores itGenerationFewer chunks, better ordering, prompt

Retrieval fixes

  • Query rewriting. Users write "what about the second one?" or use slang. An LLM rewrites the question into a standalone, document-style query, which is essential for follow-up turns.
  • HyDE. Generate a hypothetical answer and embed that; it often sits closer to real answer passages. Useful in sparse domains.
  • Hybrid search. Exact identifiers, error codes and SKUs are what embeddings blur; BM25 catches them.
  • Filters. A too-strict metadata filter (wrong department, wrong year) silently removes the right document and looks exactly like an embedding failure.

Chunking fixes

When the answer straddles a chunk boundary, add some overlap, use parent-document retrieval, or use contextual retrieval: before embedding, prepend each chunk with a short generated sentence describing where it sits in its document ("From the 2025 travel policy, section on international per-diem…"), so an isolated chunk still carries its context.

Python
def contextualise(doc_title: str, section: str, chunk: str) -> str:    summary = small_llm(f"In one sentence, say what this passage covers within '{doc_title}', "                        f"section '{section}'.\n\n{chunk}")    return f"{summary}\n\n{chunk}"          # embed and BM25-index this text; show the original

A real-life example

Scenario, numbers made up. An IT helpdesk bot often says "I couldn't find that" for questions like "VPN error 809 on my laptop". The team labels 150 failing queries. Recall@100 is 72%; of the misses, most contain an error code or a product name.

They add BM25 to form a hybrid search, query rewriting for follow-up turns, and contextual headers on chunks from long troubleshooting guides. Recall@100 rises to 95% and recall@5 from 48% to 81%. They also find that 11 failures came from a filter that restricted searches to the user's department, although VPN guides were stored under "IT-global".

Follow-up questions to expect

  • "Why not just use a bigger embedding model?" — It may help, but if the failures are error codes, filters or chunk boundaries, a bigger model will not fix them. Diagnose first.
  • "When does HyDE hurt?" — When the model's hypothetical answer is confidently wrong or in a different style from your documents, it can pull retrieval off course. Test it on the golden set.
  • "How do you label the gold chunk quickly?" — Take failed queries, retrieve the top 100, and have a domain expert pick the right chunk, or write the answer and search for it.