LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

How can you detect hallucinations in model outputs?


Five claims in one leave-policy answer26days8 carryforwardneedsapprovalpaidon exitupdatedApril01234not inthe policyFaithfulness = 4 of 5 = 0.8.
The average looks fine, but the one unsupported claim is exactly the one an employee would act on.

What you need to know

Case 1: you have a context (the easy case)

  1. Decompose — split the answer into atomic claims, one fact each.
  2. Verify — for each claim, ask an NLI model or LLM judge: is it supported by the context?
  3. Score — faithfulness = supported claims / total claims.
  4. Flag — any unsupported claim, not just a low average; one invented number can be the whole problem.

Worked example. The HR assistant answers: "You get 26 days of annual leave. Up to 8 can be carried forward. Carry-forward needs manager approval. Unused leave is paid out on exit. The policy was updated in April." The policy text supports the first three claims and the last, but says nothing about payment on exit. Faithfulness = 4 / 5 = 0.8, and the flagged claim is exactly the one an employee might act on.

RAGAS (Faithfulness) and DeepEval (FaithfulnessMetric) implement this claim-based pattern with an LLM. Small trained classifiers such as Vectara's HHEM do a similar check more cheaply.

Case 2: no context

  • Sampling consistency — generate several answers; facts the model really knows stay the same, fabrications vary (SelfCheckGPT).
  • Token log-probabilities — low probabilities on names and numbers correlate with fabrication. A weak signal on its own, useful as one feature.
  • Search-based verification — retrieve evidence for each claim from the web or your knowledge base, then check entailment.

Measure the detector

A hallucination detector is a classifier. Label 200 answers by hand and compute its precision (flagged answers that really hallucinated) and recall (hallucinations it caught). A detector with 60% recall that everyone treats as complete is dangerous.

A real-life example

The HR assistant handles 8,000 questions a day. The team runs two layers:

  • All traffic — a deterministic check that every number and date in the answer appears in the retrieved chunks. It flags about 4% of answers.
  • 5% sample plus all flagged answers — claim-level faithfulness with an LLM judge.

On 200 human-labelled answers, the number check alone has recall 0.52 (it misses invented conditions like "only for permanent staff") and precision 0.81. The claim-level judge has recall 0.88 and precision 0.84. Combined, the dashboard shows a daily unsupported-claim rate of about 3%, and every flagged answer with an unsupported number goes to a review queue.

Follow-up questions to expect

  • "Is a faithful answer always correct?" — No. If the retriever fetched an outdated policy, a perfectly faithful answer repeats the outdated fact; faithfulness measures grounding, not truth.
  • "What threshold would you alert on?" — Track the rate of answers with any unsupported claim, not average faithfulness; alert on a change from the baseline, sized by its confidence interval.
  • "How would you reduce hallucination once detected?" — Improve retrieval, instruct and allow the model to say "not in the policy", require quotes, and block or regenerate answers with unsupported numbers.