Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your live RAG assistant gets about a third of answers wrong even though the documents contain them. How do you debug it?
What you need to know
A RAG system has two halves that fail in different ways. Retrieval finds passages. Generation writes an answer from them. If you tune the prompt when retrieval is broken, or re-chunk when the prompt is broken, you waste a sprint and nothing improves.
The split test
- Collect — 50 failing questions from logs and thumbs-down feedback.
- Label — for each, a person finds the gold chunk in the index.
- Compare — was the gold chunk in the top-k the system retrieved?
- Bucket — "not retrieved" is a retrieval problem; "retrieved but answered wrong" is a generation problem.
This needs per-request logs: the rewritten query, retrieved chunk ids with scores, and the exact prompt sent. If you don't have those, adding them is the first task, because without them you are debugging a system you cannot see.
Retrieval causes, roughly in order of frequency
| Cause | What it looks like | Fix |
|---|---|---|
| Bad chunking | The answer is split across two chunks; neither scores well | Structure-aware chunks, overlap, parent-document expansion |
| Vocabulary mismatch | User says "cancel", document says "terminate" | Hybrid BM25 plus dense search, query rewriting |
| Wrong embedding model | Domain terms (drug names, part numbers) land in odd places | Try a domain-suited model on your eval set |
| Silent metadata filter | A date or tenant filter excludes the right document | Log filters; test with filters off |
Generation causes
- Too many distractors. Ten chunks where one is relevant; the model blends them. Rerank and pass fewer.
- No permission to abstain. The prompt never says "if the context does not contain the answer, say so", so the model guesses.
- Poor ordering and layout. Chunks with no source labels, or the question buried above 8,000 tokens of context. Put the question after the context and label each chunk.
Metrics that keep it fixed
- Retrieval: recall@k (was the gold chunk in the top k?) and MRR (how high was it ranked?).
- Generation: faithfulness (is every claim supported by the retrieved text?) and answer correctness against a reference.
Track them weekly on a fixed eval set, and add every new production failure to that set.
A real-life example
Scenario (illustrative numbers). An insurance company's policy assistant answers 34% of test questions wrongly. The team pulls 50 failures and labels the gold chunks.
In 31 of the 50, the gold chunk was never retrieved. Most of those questions use everyday words ("my bike was stolen") while the policy says "theft of insured two-wheeler". In the other 19, the chunk was retrieved but the answer was wrong; the prompt sent 12 chunks and never allowed "not covered in the documents".
They add BM25 alongside dense search plus a query-rewrite step, which lifts recall@5 from 0.58 to 0.84. They cut the context to the top 5 reranked chunks and add an abstain instruction. The wrong-answer rate falls to 11%, and the remaining failures are mostly questions the documents really do not answer, which now get "I couldn't find this in your policy".
Follow-up questions to expect
- "What if the gold chunk doesn't exist at all?" — Then it is a content gap, not a system bug. Log those questions and send them to the document owners.
- "How do you label 50 gold chunks quickly?" — Search the index by keyword yourself; it usually takes a few minutes per question. A domain expert helps for tricky ones.
- "Can an LLM judge do the labelling?" — It can suggest candidates, but check a sample by hand; a wrong gold label corrupts every metric downstream.