Course Content
LLM Evaluation
6 sections · 50 lessons
What types of prompts are most likely to trigger hallucinations?
What you need to know
A model produces the most likely continuation. When it does not know the answer, the most likely continuation is still a confident-looking answer — because training text is full of confident answers and rarely says "I don't know".
High-risk prompt types
| Trigger | Example | Why it fails |
|---|---|---|
| Long-tail specifics | "What year did this small NBFC get its licence?" | Rare facts are poorly memorised |
| False premise | "Why did HR discontinue the 4-day work week?" (it never existed) | The model accepts the premise and explains it |
| Citations and links | "Give three papers with DOIs" | The format is easy to copy, the content is not |
| Post-cutoff facts | "Current repo rate?" | Training data is old |
| Forced counts | "List exactly 10 benefits" | Fills slots with inventions |
| Ambiguous input | "What's the leave policy?" (which country, which grade?) | Guesses the missing context |
| Empty or off-topic retrieval | Context has no answer | Answers from memory anyway |
Why grading rewards guessing
If an eval only counts correct answers, a model that guesses on unknown questions scores higher than one that says "I don't know". Research on why language models hallucinate (including a 2025 OpenAI paper) makes exactly this point: accuracy-only grading teaches guessing. Evals must give credit, or at least no penalty, for correct abstention.
Design lessons
- Build eval sets deliberately from these triggers.
- Give the model an allowed way out: "If the policy does not cover this, say so and point to HR."
- Grade abstention: correct abstention on unanswerable items is a success, not a failure.
A real-life example
The HR assistant team writes 60 trigger questions: 20 false premises ("How do I apply for the sabbatical programme?" — there is none), 20 questions whose answer is not in any document, and 20 asking for exact figures from rarely used policies.
The first version answers 14 of 20 false-premise questions with invented procedures, including a made-up sabbatical form number. After adding "If the documents do not mention it, say it is not covered and suggest contacting HR" and a check that rejects answers with no supporting chunk, false-premise hallucinations drop to 2 of 20. Over-refusal on 200 normal questions rises from 1% to 3%, which the team accepts and tracks.
Follow-up questions to expect
- "How do you test false premises at scale?" — Take real questions and change one detail to something that does not exist (a policy, product or date); grade whether the answer corrects the premise.
- "Does RAG remove hallucination?" — It reduces it for covered topics, but the model can still ignore context, misread it, or answer from memory when retrieval misses.
- "What about reasoning models?" — They can still hallucinate facts; better reasoning does not create missing knowledge. Test them on the same triggers.