Course Content
LLM Evaluation
6 sections · 50 lessons
What makes evaluating LLMs different from traditional ML models?
What you need to know
A traditional model, such as a fraud classifier on tabular data, maps an input to one of a few labels. You compare with the true label and count. LLM evaluation breaks nearly every assumption behind that:
| Traditional ML | LLM systems |
|---|---|
| One correct label | Many correct answers, worded differently |
| One main metric | Many dimensions: correctness, faithfulness, tone, safety, format, cost, latency |
| Deterministic at inference | Sampling gives different outputs per run |
| Private held-out test set | Public benchmarks likely seen in pretraining |
| The model is the system | Prompts, retrieval, tools and orchestration all change the score |
| Metric is a formula | The grader is often another LLM that needs its own validation |
What changes in practice
- Graded, not binary, scoring. You use rubrics and partial credit, or pairwise preference.
- Repeats and intervals. Run each case more than once and report confidence intervals, because a 2-point difference may be sampling noise.
- Own the test set. Leaderboard numbers help shortlist models; your own eval on your data decides.
- Evaluate the whole pipeline, and each component separately.
- Validate the evaluator against human labels.
Where classic metrics still apply
When the LLM task is really classification or extraction — a complaint label, an invoice amount, a yes/no — use classic metrics: accuracy, precision, recall, F1, confusion matrices. Many production LLM features are closed tasks in disguise, and that is good news for evaluation.
A real-life example
The bank has two LLM features for complaints. The classifier outputs one of six labels. Evaluation is classic: on 400 labelled complaints, fraud precision is 0.80 and recall is 0.72, and a confusion matrix shows upi and fraud are confused most.
The reply drafter writes a response to the customer. There is no single right reply. The team evaluates it with deterministic checks (ticket number present, no promised refund amount, under 120 words), a judge rubric ("states the next step", "does not admit liability", "polite"), and weekly review by two customer-service leads. They also run each case three times, because one draft in twenty promises a timeline the bank never commits to — something a single run would miss.
Follow-up questions to expect
- "Is accuracy ever enough for an LLM?" — Yes, when the output is a closed label or an extracted value; turn the task into a constrained format and use classic metrics.
- "Why not trust MMLU-style leaderboard scores?" — They are measured on general tasks, may be contaminated, and are run with harnesses that differ from yours; they don't predict performance on your domain.
- "How does non-determinism change the process?" — Run repeats, report means with confidence intervals, and set regression thresholds from measured run-to-run noise.