Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Scenario – 1: Partial Answer Accuracy


Giving each part of the question its own evidenceSplit intosub-questionsRetrieve top5 for eachMerge andrerank the unionAnswer eachsub-questionCoverage check,or name the gap
One shared top-5 spends every slot on the dominant topic, so the second question never reaches the model at all.

What you need to know

The scenario: users say answers are "half right": the part answered is correct, but part of the question is ignored.

Why one retrieval starves the second question

"What is the refund window for electronics, and can I get cash instead of store credit?" produces one query embedding. That vector sits closest to whichever topic dominates the wording, so the top 5 chunks are all about refund windows. Nothing about cash refunds reaches the model, and the model answers what it can see.

The fix

  1. Detect — a cheap classifier or a small LLM call returns a list of sub-questions (a list of one for simple questions).
  2. Retrieve per sub-question — top 5 each, not top 5 shared.
  3. Merge — deduplicate, rerank the union, and keep sources labelled by sub-question.
  4. Generate once — the prompt lists the sub-questions and asks for an answer to each.
  5. Check coverage — confirm each sub-question received an answer; re-retrieve for a gap, or say it clearly.
Python
async def answer(question: str):    subqs = await decompose(question)                  # e.g. ["refund window for electronics", "cash refund allowed?"]    per_q = await asyncio.gather(*(retrieve(q, k=5) for q in subqs))    context = rerank_union(question, per_q, keep=8)    reply = await generate(question, subqs, context)    missing = [q for q in subqs if not covers(reply, q)]  # small judge call per sub-question    if missing:        reply += "\n\nI couldn't find documented information on: " + "; ".join(missing)    return reply

An honest "I couldn't find X" is better than silent omission, especially in support, where users assume the answer is complete.

Other causes to rule out

CauseSignFix
Answer split across chunksThe gold chunk holds only half the answerParent-document retrieval: match small chunks, return the whole section
Prompt rewards brevityModel stops after the first pointAsk for one answer per sub-question
Context too smallEvidence retrieved but cut by the budgetPer-intent budget, compression

Measure it

Plain accuracy hides this failure. Score sub-questions answered divided by sub-questions asked with an LLM judge on a golden set that includes multi-part questions.

A real-life example

Scenario, numbers made up. An airline's support bot handles questions like "Can I change my flight date, and will I get a refund for the seat I paid for?" A review finds that 38% of multi-part questions get an answer to only one part.

The team adds decomposition with a small model, five chunks per sub-question, and a coverage check. On a 200-question golden set with 80 multi-part questions, sub-question coverage rises from 64% to 93%. Latency rises by about 300 ms for multi-part questions only, because single questions skip decomposition. Repeat contacts about "the second part of my question" fall noticeably.

Follow-up questions to expect

  • "Doesn't decomposition add cost?" — One small model call and a few extra retrievals, only for questions detected as multi-part; far cheaper than a repeat contact.
  • "What if sub-questions depend on each other?" — Answer them in sequence, feeding the first answer into the second retrieval, rather than in parallel.
  • "How do you check coverage cheaply?" — A small model judges each sub-question against the answer with a yes/no rubric.