Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Scenario – 6: Grounding Failure and Hallucinations
What you need to know
The scenario: the right chunks are in the prompt, yet answers contain details that are not in any of them.
Why models stray from good context
| Cause | What happens | Fix |
|---|---|---|
| Too much context | With 15 chunks, the key passage in the middle gets less attention ("lost in the middle") | Rerank to 3–5; strongest evidence first |
| Partial relevance | The model fills gaps from its training data | Abstain below a reranker threshold |
| Conflicting sources | The model blends them into one answer | Tell it to show both positions with dates |
| No permission to refuse | "I don't know" is never produced | Make refusal an explicit, allowed output |
Prompt and format
Number the sources and require a citation per sentence. Forcing the model to name its evidence tends to reduce unsupported statements, and it makes failures auditable.
Answer only from the numbered sources below.After each sentence, cite the source IDs in brackets, e.g. [S2].If the sources do not contain the answer, reply exactly: "The documents I have don't cover this."If sources disagree, state both positions with their dates.Sources:[S1] (Travel policy v7, 2026-03) ...[S2] (Travel policy v6, 2025-01) ...Verify, don't trust
- Split — break the answer into single claims.
- Check — an NLI model tests each claim against its cited chunks: entailed, neutral or contradicted.
- Act — strip or regenerate unsupported claims; hard-block unsupported numbers and policy statements.
- Audit — a human reviews about 100 answers a week to keep the automatic checker honest.
A small NLI model adds only tens of milliseconds per answer on a GPU, so this can run on every response. Track faithfulness as a release gate on the golden set.
A real-life example
Scenario, numbers made up. A corporate travel assistant retrieves the right policy chunks, but answers sometimes say "business class is allowed on flights over 6 hours" when the policy says 8. An audit finds 7% of answers contain at least one unsupported claim, most with 12–15 chunks in the prompt.
The team reranks to 4 chunks, adds per-sentence citations and an explicit "documents don't cover this" reply, and runs an NLI check on every answer. Unsupported claims fall to about 1%, and the remaining ones are mostly caught and stripped before display. The weekly audit also finds that two old policy versions were still indexed, which is why some answers blended the 6- and 8-hour rules.
Follow-up questions to expect
- "Why not use a larger context model?" — More room does not fix attention to the middle or gap-filling; fewer, better chunks usually improve faithfulness.
- "Do citations prove the answer is grounded?" — No; models sometimes cite a nearby but wrong source. That is why each citation is verified.
- "How do you handle outdated documents?" — Store effective dates, filter superseded versions, and show dates when sources conflict.