Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Scenario – 1: Partial Answer Accuracy
What you need to know
The scenario: users say answers are "half right": the part answered is correct, but part of the question is ignored.
Why one retrieval starves the second question
"What is the refund window for electronics, and can I get cash instead of store credit?" produces one query embedding. That vector sits closest to whichever topic dominates the wording, so the top 5 chunks are all about refund windows. Nothing about cash refunds reaches the model, and the model answers what it can see.
The fix
- Detect — a cheap classifier or a small LLM call returns a list of sub-questions (a list of one for simple questions).
- Retrieve per sub-question — top 5 each, not top 5 shared.
- Merge — deduplicate, rerank the union, and keep sources labelled by sub-question.
- Generate once — the prompt lists the sub-questions and asks for an answer to each.
- Check coverage — confirm each sub-question received an answer; re-retrieve for a gap, or say it clearly.
1async def answer(question: str):2 subqs = await decompose(question) # e.g. ["refund window for electronics", "cash refund allowed?"]3 per_q = await asyncio.gather(*(retrieve(q, k=5) for q in subqs))4 context = rerank_union(question, per_q, keep=8)5 reply = await generate(question, subqs, context)6 missing = [q for q in subqs if not covers(reply, q)] # small judge call per sub-question7 if missing:8 reply += "\n\nI couldn't find documented information on: " + "; ".join(missing)9 return replyAn honest "I couldn't find X" is better than silent omission, especially in support, where users assume the answer is complete.
Other causes to rule out
| Cause | Sign | Fix |
|---|---|---|
| Answer split across chunks | The gold chunk holds only half the answer | Parent-document retrieval: match small chunks, return the whole section |
| Prompt rewards brevity | Model stops after the first point | Ask for one answer per sub-question |
| Context too small | Evidence retrieved but cut by the budget | Per-intent budget, compression |
Measure it
Plain accuracy hides this failure. Score sub-questions answered divided by sub-questions asked with an LLM judge on a golden set that includes multi-part questions.
A real-life example
Scenario, numbers made up. An airline's support bot handles questions like "Can I change my flight date, and will I get a refund for the seat I paid for?" A review finds that 38% of multi-part questions get an answer to only one part.
The team adds decomposition with a small model, five chunks per sub-question, and a coverage check. On a 200-question golden set with 80 multi-part questions, sub-question coverage rises from 64% to 93%. Latency rises by about 300 ms for multi-part questions only, because single questions skip decomposition. Repeat contacts about "the second part of my question" fall noticeably.
Follow-up questions to expect
- "Doesn't decomposition add cost?" — One small model call and a few extra retrievals, only for questions detected as multi-part; far cheaper than a repeat contact.
- "What if sub-questions depend on each other?" — Answer them in sequence, feeding the first answer into the second retrieval, rather than in parallel.
- "How do you check coverage cheaply?" — A small model judges each sub-question against the answer with a yes/no rubric.