Course Content
RAG Systems
12 sections · 66 lessons
How do you detect hallucinations in a RAG pipeline?
What you need to know
Method 1: claim-level faithfulness
- Split the answer into atomic claims, each one simple fact.
- For each claim, ask a judge: "Can this be inferred from the context? yes or no".
- Score = yes count ÷ claim count.
A natural language inference (NLI) model can do step 2 more cheaply. It is a small classifier that labels a pair (context, claim) as entailed, contradicted or neutral.
Method 2: citation checks
Require [n] citations, then check in code that every sentence has one and every cited id was actually retrieved. It costs nothing and catches invented sources. It does not catch a claim that cites a real chunk but misstates it; that needs method 1.
Method 3: consistency across samples
Generate the answer 3 times with sampling on. Claims that change between samples are often unsupported. This is the idea behind SelfCheckGPT. It costs several extra calls, so it is used offline or on high-risk answers.
Method 4: reference comparison
On the golden set, compare with the reference answer. This catches the case the other methods miss: an answer that is faithful to a wrong or outdated chunk.
In production
| Check | Cost | Run on |
|---|---|---|
| No-context guard, citation check | Free | Every request |
| NLI faithfulness | Small model call | Every request, if latency allows |
| LLM-judge faithfulness | One extra LLM call | 1 to 5% sample, plus all high-risk answers |
| Human review | Expensive | Low-score and flagged cases |
Store the score on the trace, chart it daily, and alert when the average drops.
A real-life example
A bank's FAQ bot answers "What do I get with the Platinum card?" The retrieved context says: annual fee Rs 2,999; fee waived if you spend Rs 3 lakh in a year; 8 free airport lounge visits a year. The answer:
1. The annual fee is Rs 2,999 [1]. -> supported2. It is waived if you spend Rs 3 lakh in a year [1]. -> supported3. You get unlimited airport lounge access [2]. -> NOT supported (context says 8)4. You also get 5% cashback on fuel. -> NOT supported (not in context)Faithfulness = 2 ÷ 4 = 0.5. Notice that the citation check catches claim 4 (no citation) but not claim 3, because [2] is a real chunk. Only the claim-level judge catches the "unlimited" error. The bank routes any answer with faithfulness below 1.0 on fee or benefit questions to a fallback: "Please see the Platinum card page for full benefits," with a link.
Follow-up questions to expect
- "Can the judge itself hallucinate?" — Yes. Check it against 50 to 100 human-labelled answers, keep the question narrow (one claim, yes or no), and prefer a different model family from the generator.
- "Is a faithful answer always correct?" — No. If the retrieved chunk is outdated, a perfectly faithful answer is still wrong. That is why you also compare with references and keep the index fresh.
- "How do you reduce the judge's cost?" — Use an NLI model or a small judge for most traffic, and the large judge only on samples and high-risk topics.