LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

What is SelfCheckGPT, and how does it work without external sources?


Ten samples: how fast does it charge to 50%?25min30minnone45min25minnone30minnone40minnone0123456789matchesoriginalmatchesoriginalThe 5,000 mAh battery appeared in 10 of 10 samples; the charging claim in 2 of 10.
Known facts repeat across samples and invented ones scatter, which is the whole signal SelfCheckGPT uses.

What you need to know

The procedure

  1. Original answer — generate the response you want to check.
  2. Extra samples — generate N more responses to the same prompt with sampling turned on.
  3. Score each sentence — how strongly do the samples support this sentence of the original?
  4. Flag — sentences with low support are likely fabricated.

Scoring variants

The paper tested several ways to measure support: n-gram overlap, BERTScore, question answering (ask questions about the sentence, answer them from each sample, compare), an NLI model, and directly prompting an LLM "Is this sentence supported by this passage? Yes or no". The NLI and prompt variants performed best. There is an open-source selfcheckgpt Python package with these variants.

Worked example

A product-description generator is asked about a phone whose spec sheet only lists the battery (5,000 mAh) and screen size. The original says: "It has a 5,000 mAh battery. It charges to 50% in 25 minutes." In 10 extra samples:

SentenceSamples that agreeSupport
5,000 mAh battery10 of 101.0
50% in 25 minutes2 of 10 (others say 30 minutes, 45 minutes, or nothing)0.2

The charging claim is flagged. It was invented — nothing in the input supported it.

Strengths and limits

  • Works anywhere — any API, no knowledge base, no reference answer.
  • Costly — N + 1 generations per item, so use it offline or on sampled traffic.
  • Detects inconsistency, not falsehood — a model that consistently repeats a wrong fact it learned passes every check.
  • Needs real variety — at temperature 0, samples are near-copies and everything looks consistent.

A real-life example

The e-commerce team has 40,000 older product pages generated before they added spec-grounding checks. They cannot manually review them all, and many have no stored spec sheet. They run a SelfCheckGPT-style pass with 5 samples per page on a 2,000-page sample, using an LLM prompt to score support per sentence.

Sentences with support below 0.4 are flagged on 7% of pages. A copy editor reviews 100 flagged sentences: 71 are real fabrications (mostly charging speeds, water resistance ratings and "made in" claims). She also reviews 100 unflagged sentences and finds 6 errors — all consistent, plausible-sounding claims such as a common but wrong warranty length. The team uses the detector to prioritise the clean-up and knows its blind spot.

Follow-up questions to expect

  • "How many samples do you need?" — The paper used 20; in practice 5 to 10 gives most of the signal. Measure detection quality on labelled data as you reduce N.
  • "Why would consistent errors happen?" — If the model learned a wrong fact in training, it believes it; every sample repeats it. Only external evidence catches this.
  • "How is this related to self-consistency prompting?" — Same mechanism, different use: self-consistency votes to choose an answer; SelfCheckGPT uses disagreement to flag likely fabrication.