Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Scenario – 6: Grounding Failure and Hallucinations


What you need to know

The scenario: the right chunks are in the prompt, yet answers contain details that are not in any of them.

Why models stray from good context

CauseWhat happensFix
Too much contextWith 15 chunks, the key passage in the middle gets less attention ("lost in the middle")Rerank to 3–5; strongest evidence first
Partial relevanceThe model fills gaps from its training dataAbstain below a reranker threshold
Conflicting sourcesThe model blends them into one answerTell it to show both positions with dates
No permission to refuse"I don't know" is never producedMake refusal an explicit, allowed output

Prompt and format

Number the sources and require a citation per sentence. Forcing the model to name its evidence tends to reduce unsupported statements, and it makes failures auditable.

Text
Answer only from the numbered sources below.After each sentence, cite the source IDs in brackets, e.g. [S2].If the sources do not contain the answer, reply exactly: "The documents I have don't cover this."If sources disagree, state both positions with their dates.Sources:[S1] (Travel policy v7, 2026-03) ...[S2] (Travel policy v6, 2025-01) ...

Verify, don't trust

  1. Split — break the answer into single claims.
  2. Check — an NLI model tests each claim against its cited chunks: entailed, neutral or contradicted.
  3. Act — strip or regenerate unsupported claims; hard-block unsupported numbers and policy statements.
  4. Audit — a human reviews about 100 answers a week to keep the automatic checker honest.

A small NLI model adds only tens of milliseconds per answer on a GPU, so this can run on every response. Track faithfulness as a release gate on the golden set.

A real-life example

Scenario, numbers made up. A corporate travel assistant retrieves the right policy chunks, but answers sometimes say "business class is allowed on flights over 6 hours" when the policy says 8. An audit finds 7% of answers contain at least one unsupported claim, most with 12–15 chunks in the prompt.

The team reranks to 4 chunks, adds per-sentence citations and an explicit "documents don't cover this" reply, and runs an NLI check on every answer. Unsupported claims fall to about 1%, and the remaining ones are mostly caught and stripped before display. The weekly audit also finds that two old policy versions were still indexed, which is why some answers blended the 6- and 8-hour rules.

Follow-up questions to expect

  • "Why not use a larger context model?" — More room does not fix attention to the middle or gap-filling; fewer, better chunks usually improve faithfulness.
  • "Do citations prove the answer is grounded?" — No; models sometimes cite a nearby but wrong source. That is why each citation is verified.
  • "How do you handle outdated documents?" — Store effective dates, filter superseded versions, and show dates when sources conflict.