Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your RAG chatbot answers payment policy questions. Users report the bot sometimes makes up refund rules that don't exist in any document. How do you detect and prevent hallucinations at the response layer — before the user sees them?
What you need to know
Where invented rules come from
When retrieval returns nothing useful, the model still tries to be helpful and fills the gap with a plausible rule: "Refunds are processed within 7 days". This is the most common kind of RAG hallucination in policy bots, and it is fixable because the facts are narrow and checkable.
Layer 1: prevent
- Abstain below a calibrated reranker threshold — never generate from junk context.
- Instruct the model to answer only from context, cite a chunk ID for every rule, and say "I don't have that documented" otherwise.
- Low temperature.
- For high-risk values (refund windows, fees, limits), fetch them from a structured table and insert them, instead of letting the model read them from prose.
Layer 2: verify after generation
1from sentence_transformers import CrossEncoder23nli = CrossEncoder("cross-encoder/nli-deberta-v3-base")4LABELS = ["contradiction", "entailment", "neutral"] # this model's output order56def check(answer: str, chunks: dict[str, str]):7 claims = split_claims(answer) # a cheap model or rules split into claims8 pairs = [(chunks.get(c.cited_id, ""), c.text) for c in claims]9 scores = nli.predict(pairs) # one row of 3 scores per claim10 return [(c.text, LABELS[s.argmax()]) for c, s in zip(claims, scores)11 if LABELS[s.argmax()] != "entailment"]A small NLI model runs quickly and cheaply, so it can check every answer. Send only the flagged claims to an LLM judge for a second look. Treat unsupported numbers and unsupported rule statements as hard blocks.
Layer 3: gate
- Strip — remove the unsupported sentence if the rest still answers the question.
- Regenerate once — with a stricter prompt that names the failed claim.
- Hand over — show the relevant policy document and offer a human agent.
| Metric | Why |
|---|---|
| Groundedness rate | Share of answers where every claim is supported |
| Block rate | How often the gate fires |
| False-block rate | Good answers wrongly blocked; too high and the bot is useless |
Check the verifier itself against about 100 human-audited answers each week. An unchecked verifier is just a second model that can be wrong.
A real-life example
Scenario, numbers made up. A payments app's support bot is caught telling users that "UPI refunds above ₹10,000 need a video KYC" — a rule that exists nowhere. An audit of 500 answers finds 4% contain at least one unsupported rule, almost all from questions where retrieval found nothing relevant.
The team adds an abstain threshold, a policy table for 30 key values, and an NLI check on every answer. Unsupported rules in the next audit fall to 0.3%. The first version of the checker blocked 9% of good answers because it treated paraphrases as "neutral"; tuning the claim splitter and allowing the LLM judge to overrule NLI on paraphrase cases brings false blocks down to 2%.
Follow-up questions to expect
- "Isn't the verifier also an LLM that can hallucinate?" — Yes, which is why NLI models are small and narrow, flagged claims get a second check, and the verifier is audited against humans weekly.
- "Does this add latency?" — Some. Run the NLI check while the answer streams to a buffer, or verify before showing for high-risk intents only.
- "What about answers that are correct but not in the documents?" — In a policy bot, those are still blocked; if it is not documented, it is not policy.