Course Content
LangChain Mastery
7 sections · 109 lessons
Write a function to detect and handle LLM hallucination in LangChain.
What you need to know
Layers of defence
- Prevent — prompt: "answer only from the context; if it is missing, say you don't know"; require citations; keep temperature low for factual tasks.
- Cheap checks — no documents above the score threshold → don't generate. Cited source ids must exist in the context. Numbers in the answer must appear in the tool output.
- Judge — a second model call checks each claim against the context.
- Act — regenerate once, fall back to "I don't know", or route to a human.
The groundedness check
1from pydantic import BaseModel, Field23class Grounding(BaseModel):4 supported: bool = Field(description="True only if every claim is supported by the context")5 unsupported_claims: list[str] = Field(default_factory=list)67judge = judge_llm.with_structured_output(Grounding)89def check_grounded(answer: str, docs) -> Grounding:10 context = "\n\n".join(d.page_content for d in docs)11 return judge.invoke(12 "You are checking an answer against its sources. For each claim in the "13 "ANSWER, decide whether the CONTEXT supports it. List unsupported claims.\n\n"14 f"CONTEXT:\n{context}\n\nANSWER:\n{answer}")1516def answer_safely(question: str) -> str:17 docs = retriever.invoke(question)18 if not docs:19 return "I couldn't find this in our documents."20 answer = rag_chain.invoke(question)21 verdict = check_grounded(answer, docs)22 if verdict.supported:23 return answer24 log.warning("ungrounded", extra={"claims": verdict.unsupported_claims})25 return "I'm not certain about this. I've passed your question to the team."- The judge sees only the context and the answer, so it checks support, not general truth.
- Use a capable judge model; a weak one misses subtle errors.
- The judge can be wrong too. Treat its verdict as a signal for action and monitoring, not as proof.
Cost
The judge adds one model call per answer. Common choices: run it on every answer for high-risk domains (finance, health, HR decisions), sample 5–10% of traffic elsewhere to monitor the hallucination rate, or run it on answers that contain numbers or dates.
A real-life example
A finance assistant answers "What was Reliance's closing price yesterday, and how did it move this week?" using a stock-price tool. In testing it once said "up 4.2% this week" when the tool data showed 2.4%.
The team added a numeric check: every number in the answer must appear in the tool output or be computable from it. The wrong figure failed the check, and the assistant regenerated with "quote numbers exactly from the tool result". They also run a groundedness judge on a 10% sample of production traces in LangSmith. The weekly dashboard shows the ungrounded rate — about 1.5% — and each flagged trace goes into the evaluation dataset.
Follow-up questions to expect
- "Can you detect hallucination without sources?" — Only weakly — for example by sampling several answers and checking whether they agree. Grounding against sources is far more reliable.
- "How do you reduce it rather than detect it?" — Better retrieval, a grounded prompt that allows "I don't know", citations, structured output, and tools for facts like prices and dates.
- "How do you measure it over time?" — A faithfulness metric on a fixed dataset in CI, plus sampled judging of production traces.