LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

What makes evaluating LLMs different from traditional ML models?


What you need to know

A traditional model, such as a fraud classifier on tabular data, maps an input to one of a few labels. You compare with the true label and count. LLM evaluation breaks nearly every assumption behind that:

Traditional MLLLM systems
One correct labelMany correct answers, worded differently
One main metricMany dimensions: correctness, faithfulness, tone, safety, format, cost, latency
Deterministic at inferenceSampling gives different outputs per run
Private held-out test setPublic benchmarks likely seen in pretraining
The model is the systemPrompts, retrieval, tools and orchestration all change the score
Metric is a formulaThe grader is often another LLM that needs its own validation

What changes in practice

  • Graded, not binary, scoring. You use rubrics and partial credit, or pairwise preference.
  • Repeats and intervals. Run each case more than once and report confidence intervals, because a 2-point difference may be sampling noise.
  • Own the test set. Leaderboard numbers help shortlist models; your own eval on your data decides.
  • Evaluate the whole pipeline, and each component separately.
  • Validate the evaluator against human labels.

Where classic metrics still apply

When the LLM task is really classification or extraction — a complaint label, an invoice amount, a yes/no — use classic metrics: accuracy, precision, recall, F1, confusion matrices. Many production LLM features are closed tasks in disguise, and that is good news for evaluation.

A real-life example

The bank has two LLM features for complaints. The classifier outputs one of six labels. Evaluation is classic: on 400 labelled complaints, fraud precision is 0.80 and recall is 0.72, and a confusion matrix shows upi and fraud are confused most.

The reply drafter writes a response to the customer. There is no single right reply. The team evaluates it with deterministic checks (ticket number present, no promised refund amount, under 120 words), a judge rubric ("states the next step", "does not admit liability", "polite"), and weekly review by two customer-service leads. They also run each case three times, because one draft in twenty promises a timeline the bank never commits to — something a single run would miss.

Follow-up questions to expect

  • "Is accuracy ever enough for an LLM?" — Yes, when the output is a closed label or an extracted value; turn the task into a constrained format and use classic metrics.
  • "Why not trust MMLU-style leaderboard scores?" — They are measured on general tasks, may be contaminated, and are run with harnesses that differ from yours; they don't predict performance on your domain.
  • "How does non-determinism change the process?" — Run repeats, report means with confidence intervals, and set regression thresholds from measured run-to-run noise.