LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

When is self-evaluation useful, and what risks does it introduce?


What you need to know

Where it works

  • Checks against evidence — "Does each claim appear in these passages? Quote the passage." The model is comparing with text, not trusting its memory.
  • Checks against a tool — run the unit tests, execute the SQL, call the validator. The tool is the judge; the model only reads the result.
  • Constraint checks — valid JSON, all fields present, within word limit. Cheap repair loops before returning.
  • Sampling consistency — generate several answers and measure agreement. Disagreement flags likely hallucinations (see SelfCheckGPT in the next section).

Where it fails

  • Unchecked self-critique — asking "are you sure?" without evidence often yields the same answer restated with confidence, or an unnecessary change from a right answer to a wrong one.
  • Calibration — stated confidence ("I'm 95% sure") is a weak predictor of correctness.
  • Self-preference — the model favours its own style and content.
  • Consistent errors — if the model learned a wrong fact, every sample repeats it.

Repair vs report

Use self-evaluation to improve outputs in the pipeline (generate, check against evidence, fix). Use an independent judge or humans to measure the result. The number shown to a stakeholder must not come from the system grading itself.

A real-life example

The code-review bot drafts a comment, then runs a self-check step: "For each claim about the code, point to the exact line in the diff. If you cannot, delete the claim." Comments mentioning variables that don't exist in the diff drop from 9% to 2% — a real improvement, because the check compares with the diff.

The team briefly also asked the bot to rate its own comment's usefulness from 1 to 5 and put that average in their weekly report: 4.4 out of 5. The independent strong-model judge, calibrated against engineers, rated the same comments 61% useful. They removed the self-rating from the report and kept only the grounded self-check in the pipeline.

Follow-up questions to expect

  • "Does self-reflection improve accuracy on reasoning tasks?" — It helps most when there is external feedback (test results, tool errors); without feedback, gains are small and it can change correct answers to wrong ones.
  • "How is self-consistency different from self-evaluation?" — Self-consistency samples several independent answers and votes; the model never grades itself. The agreement rate is a signal, not a self-assessment.
  • "Can a model's token probabilities be a self-check?" — Low probability on key tokens correlates with errors and is a useful feature, but it is noisy and not available from every API.