LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

What types of prompts are most likely to trigger hallucinations?


What you need to know

A model produces the most likely continuation. When it does not know the answer, the most likely continuation is still a confident-looking answer — because training text is full of confident answers and rarely says "I don't know".

High-risk prompt types

TriggerExampleWhy it fails
Long-tail specifics"What year did this small NBFC get its licence?"Rare facts are poorly memorised
False premise"Why did HR discontinue the 4-day work week?" (it never existed)The model accepts the premise and explains it
Citations and links"Give three papers with DOIs"The format is easy to copy, the content is not
Post-cutoff facts"Current repo rate?"Training data is old
Forced counts"List exactly 10 benefits"Fills slots with inventions
Ambiguous input"What's the leave policy?" (which country, which grade?)Guesses the missing context
Empty or off-topic retrievalContext has no answerAnswers from memory anyway

Why grading rewards guessing

If an eval only counts correct answers, a model that guesses on unknown questions scores higher than one that says "I don't know". Research on why language models hallucinate (including a 2025 OpenAI paper) makes exactly this point: accuracy-only grading teaches guessing. Evals must give credit, or at least no penalty, for correct abstention.

Design lessons

  • Build eval sets deliberately from these triggers.
  • Give the model an allowed way out: "If the policy does not cover this, say so and point to HR."
  • Grade abstention: correct abstention on unanswerable items is a success, not a failure.

A real-life example

The HR assistant team writes 60 trigger questions: 20 false premises ("How do I apply for the sabbatical programme?" — there is none), 20 questions whose answer is not in any document, and 20 asking for exact figures from rarely used policies.

The first version answers 14 of 20 false-premise questions with invented procedures, including a made-up sabbatical form number. After adding "If the documents do not mention it, say it is not covered and suggest contacting HR" and a check that rejects answers with no supporting chunk, false-premise hallucinations drop to 2 of 20. Over-refusal on 200 normal questions rises from 1% to 3%, which the team accepts and tracks.

Follow-up questions to expect

  • "How do you test false premises at scale?" — Take real questions and change one detail to something that does not exist (a policy, product or date); grade whether the answer corrects the premise.
  • "Does RAG remove hallucination?" — It reduces it for covered topics, but the model can still ignore context, misread it, or answer from memory when retrieval misses.
  • "What about reasoning models?" — They can still hallucinate facts; better reasoning does not create missing knowledge. Test them on the same triggers.