Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your RAG cites the correct source document, but the generated answer contradicts what the document says. How do you debug and fix faithfulness failures in production?


Right source, wrong answer: four causesUnfaithful answerOld and newpolicy both retrievedModel's prior beats the textChunk holds 80%,model invents the restEvidence lost mid-context
An instruction to use only the context fixes one cause; the other three are fixed in retrieval.

What you need to know

Reproduce the exact input first

You need the full prompt as it was sent: retrieved chunks, their order, the system prompt, the model and its version. Without that log you are debugging a different system from the one that failed. If it is not logged per request, that is the first fix.

The four causes

CauseWhat you see in the traceFix
Conflicting contextTwo chunks disagree, such as the 2024 and 2026 versions of a policyFilter to the current version; boost recency; remove superseded documents
Prior overrides contextThe document says something unusual; the model answers what is "usually" trueInstruct: answer only from the context, and say when it differs from general knowledge
Partial contextThe chunk holds 80% of the answer; the model invents the restBetter coverage: parent-chunk or whole-section expansion
Lost in a long contextThe evidence sits in the middle of 15 chunksRerank and cut to about 5 chunks

Partial context is the most common in practice. The model is trying to be helpful and completes the pattern.

A verification pass

Python
claims = llm.split_into_claims(answer)                  # short, atomic statementsresults = [judge.is_supported(claim, context) for claim in claims]unsupported = [c for c, ok in zip(claims, results) if not ok]if unsupported:    log_unfaithful(request_id, unsupported)             # log-only mode first    # later: block, regenerate, or add "I could not confirm..." to the answer

A small, cheap model checks each claim against the retrieved chunks. It costs one extra call per answer and is worth it in regulated domains like finance, insurance and health.

  1. Log-only — run the verifier on all traffic, block nothing; measure the unsupported-claim rate.
  2. Fix the biggest cause — using the traces the verifier flags.
  3. Enforce — for high-risk topics, regenerate once with the failing claim pointed out, or answer with only the supported part.

Metrics

Faithfulness on a golden set, and the online unsupported-claim rate from the verifier. Watch the verifier itself too: check a sample of its flags by hand, since a verifier that flags everything gets ignored.

A real-life example

Scenario (illustrative numbers). A health insurer's assistant tells a customer that cataract surgery has a 2-year waiting period, citing the policy wording document. The document actually says 1 year for this plan. The trace shows two retrieved chunks: this year's wording (1 year) and a 2023 version (2 years) that was never removed from the index.

The team adds a valid_to field and filters out superseded versions. They run a claim verifier in log-only mode for two weeks: 4.8% of answers contain at least one unsupported claim, and 60% of those come from partial context in waiting-period and co-pay questions. Expanding retrieval to whole policy sections and enforcing the verifier on claims-related topics brings the rate to 0.9%.

Follow-up questions to expect

  • "Can't a stronger model fix this?" — It helps with some cases, but conflicting chunks and partial context produce wrong answers from any model. Fix the inputs first.
  • "How do you check faithfulness offline?" — A golden set with reference answers, scored by a judge that you have calibrated against human labels.
  • "Doesn't verification add latency?" — One small-model call, which can run while the answer streams; for high-risk topics, hold the final sentence until it passes.