RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

How do you detect hallucinations in a RAG pipeline?


Platinum card answer: two checks, four claimspasssupportedpasssupportedpassNOT supportedFAILNOT supportedCitation checkClaim judgeFee Rs 2,999 [1]Waived at Rs 3 lakh [1]Unlimited lounge [2]5% fuel cashbackFaithfulness = 2 supported of 4 claims = 0.5.
A real citation on a wrong claim passes the cheap check, which is why claim-level faithfulness scoring is still needed.

What you need to know

Method 1: claim-level faithfulness

  1. Split the answer into atomic claims, each one simple fact.
  2. For each claim, ask a judge: "Can this be inferred from the context? yes or no".
  3. Score = yes count ÷ claim count.

A natural language inference (NLI) model can do step 2 more cheaply. It is a small classifier that labels a pair (context, claim) as entailed, contradicted or neutral.

Method 2: citation checks

Require [n] citations, then check in code that every sentence has one and every cited id was actually retrieved. It costs nothing and catches invented sources. It does not catch a claim that cites a real chunk but misstates it; that needs method 1.

Method 3: consistency across samples

Generate the answer 3 times with sampling on. Claims that change between samples are often unsupported. This is the idea behind SelfCheckGPT. It costs several extra calls, so it is used offline or on high-risk answers.

Method 4: reference comparison

On the golden set, compare with the reference answer. This catches the case the other methods miss: an answer that is faithful to a wrong or outdated chunk.

In production

CheckCostRun on
No-context guard, citation checkFreeEvery request
NLI faithfulnessSmall model callEvery request, if latency allows
LLM-judge faithfulnessOne extra LLM call1 to 5% sample, plus all high-risk answers
Human reviewExpensiveLow-score and flagged cases

Store the score on the trace, chart it daily, and alert when the average drops.

A real-life example

A bank's FAQ bot answers "What do I get with the Platinum card?" The retrieved context says: annual fee Rs 2,999; fee waived if you spend Rs 3 lakh in a year; 8 free airport lounge visits a year. The answer:

Text
1. The annual fee is Rs 2,999 [1].                          -> supported2. It is waived if you spend Rs 3 lakh in a year [1].        -> supported3. You get unlimited airport lounge access [2].             -> NOT supported (context says 8)4. You also get 5% cashback on fuel.                         -> NOT supported (not in context)

Faithfulness = 2 ÷ 4 = 0.5. Notice that the citation check catches claim 4 (no citation) but not claim 3, because [2] is a real chunk. Only the claim-level judge catches the "unlimited" error. The bank routes any answer with faithfulness below 1.0 on fee or benefit questions to a fallback: "Please see the Platinum card page for full benefits," with a link.

Follow-up questions to expect

  • "Can the judge itself hallucinate?" — Yes. Check it against 50 to 100 human-labelled answers, keep the question narrow (one claim, yes or no), and prefer a different model family from the generator.
  • "Is a faithful answer always correct?" — No. If the retrieved chunk is outdated, a perfectly faithful answer is still wrong. That is why you also compare with references and keep the index fresh.
  • "How do you reduce the judge's cost?" — Use an NLI model or a small judge for most traffic, and the large judge only on samples and high-risk topics.