Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your RAG cites the correct source document, but the generated answer contradicts what the document says. How do you debug and fix faithfulness failures in production?
What you need to know
Reproduce the exact input first
You need the full prompt as it was sent: retrieved chunks, their order, the system prompt, the model and its version. Without that log you are debugging a different system from the one that failed. If it is not logged per request, that is the first fix.
The four causes
| Cause | What you see in the trace | Fix |
|---|---|---|
| Conflicting context | Two chunks disagree, such as the 2024 and 2026 versions of a policy | Filter to the current version; boost recency; remove superseded documents |
| Prior overrides context | The document says something unusual; the model answers what is "usually" true | Instruct: answer only from the context, and say when it differs from general knowledge |
| Partial context | The chunk holds 80% of the answer; the model invents the rest | Better coverage: parent-chunk or whole-section expansion |
| Lost in a long context | The evidence sits in the middle of 15 chunks | Rerank and cut to about 5 chunks |
Partial context is the most common in practice. The model is trying to be helpful and completes the pattern.
A verification pass
1claims = llm.split_into_claims(answer) # short, atomic statements2results = [judge.is_supported(claim, context) for claim in claims]3unsupported = [c for c, ok in zip(claims, results) if not ok]4if unsupported:5 log_unfaithful(request_id, unsupported) # log-only mode first6 # later: block, regenerate, or add "I could not confirm..." to the answerA small, cheap model checks each claim against the retrieved chunks. It costs one extra call per answer and is worth it in regulated domains like finance, insurance and health.
- Log-only — run the verifier on all traffic, block nothing; measure the unsupported-claim rate.
- Fix the biggest cause — using the traces the verifier flags.
- Enforce — for high-risk topics, regenerate once with the failing claim pointed out, or answer with only the supported part.
Metrics
Faithfulness on a golden set, and the online unsupported-claim rate from the verifier. Watch the verifier itself too: check a sample of its flags by hand, since a verifier that flags everything gets ignored.
A real-life example
Scenario (illustrative numbers). A health insurer's assistant tells a customer that cataract surgery has a 2-year waiting period, citing the policy wording document. The document actually says 1 year for this plan. The trace shows two retrieved chunks: this year's wording (1 year) and a 2023 version (2 years) that was never removed from the index.
The team adds a valid_to field and filters out superseded versions. They run a claim verifier in log-only mode for two weeks: 4.8% of answers contain at least one unsupported claim, and 60% of those come from partial context in waiting-period and co-pay questions. Expanding retrieval to whole policy sections and enforcing the verifier on claims-related topics brings the rate to 0.9%.
Follow-up questions to expect
- "Can't a stronger model fix this?" — It helps with some cases, but conflicting chunks and partial context produce wrong answers from any model. Fix the inputs first.
- "How do you check faithfulness offline?" — A golden set with reference answers, scored by a judge that you have calibrated against human labels.
- "Doesn't verification add latency?" — One small-model call, which can run while the answer streams; for high-risk topics, hold the final sentence until it passes.