Course Content
AI Safety & Guardrails
5 sections · 50 lessons
What are hallucinations in LLMs, and how can they be reduced?
What you need to know
Why models hallucinate
- The objective. A language model predicts the next token. When it has no reliable signal, it still produces the most plausible-sounding text, because silence is not a token it was trained to prefer.
- The incentives. Most benchmarks give 1 point for a right answer and 0 for both a wrong answer and "I don't know". Under that scoring, guessing always wins, so models learn to guess. OpenAI made this argument directly in a 2025 paper on why models hallucinate.
- Missing or stale knowledge. Facts after the training cut-off, your company's private data, and rare facts seen only a few times in training are where confident errors cluster.
- Bad context. In a RAG system the model may be faithful to a retrieved chunk that is itself wrong, outdated or about a different product.
What reduces it, in order of leverage
- Ground the answer. Retrieve evidence, put it in the prompt, tell the model to answer only from it and cite the passage. Recall becomes reading comprehension, which models are much better at.
- Allow abstention. Write "if the context does not answer the question, say so" into the prompt, and count a correct abstention as a pass in your evals. If you punish "I don't know", you train bluffing.
- Verify after generation. Check each claim against its cited passage with an entailment (NLI) model or an LLM judge. On failure, regenerate once or fall back to a refusal.
- Constrain the output. Use structured outputs with schema validation, and do calculations, dates and lookups in code. A model should never compute an EMI; a function should.
- Tune the path. Lower temperature on factual routes where the provider allows it, and use a stronger or reasoning model for high-stakes questions.
How to measure it
Keep three numbers on a fixed eval set: a groundedness (faithfulness) score for answerable questions, a citation-support rate (does the cited passage actually support the sentence), and an abstention rate on a deliberately unanswerable set. If abstention on unanswerable questions is near zero, the model is bluffing.
A real-life example
A private bank in India launches a customer chatbot. A customer asks, "Is there a penalty if I close my floating-rate home loan early?" The first version, with no retrieval, says "Yes, 2% of the outstanding amount" — a common rule at some lenders, but wrong for this bank, which charges nothing on floating-rate loans to individuals.
The team makes three changes. Retrieval pulls the bank's current loan policy page. The prompt says to answer only from that page and quote the line. A verifier checks that the quoted line appears in the page. On a 300-question eval set, unsupported claims fall from 11% to 2%, and on 50 unanswerable questions ("What will the repo rate be next quarter?") abstention rises from 8% to 90%. The remaining 2% come mostly from an outdated PDF still in the index, which is a data fix, not a prompt fix.
Follow-up questions to expect
- "Can you eliminate hallucination?" — No. You can make it rare and detectable, and design the product so a wrong answer is caught or is cheap. Any claim of "zero hallucinations" is a red flag.
- "Does a bigger model fix it?" — It usually lowers the rate on common facts, but not on your private or recent data, and a bigger model's errors are more convincing. Grounding still matters.
- "How do you detect it without ground-truth answers?" — Check faithfulness to the retrieved context with an NLI model or judge, and sample several answers: when samples disagree with each other, the model is often guessing.