LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

How do you debug poor model performance systematically?


Fraud class of the complaint classifier, 1,000 complaints722818882predicted fraudpredicted otheractually fraudnot fraudPrecision 0.80, recall 0.72, accuracy 95.4%.
Accuracy looks excellent because fraud is rare; the 28 missed cases are the ones worth reading.

What you need to know

Error analysis

  1. Collect failures — 20 to 50 from the eval or from production thumbs-down, not one anecdote.
  2. Read full traces — the final prompt after templating, retrieved chunks, tool calls and outputs.
  3. Label and count — free-text notes first, then categories.
  4. Isolate the stage — retrieval, prompt, model, tool, post-processing.
  5. Change one thing — and re-run the whole suite, not only the failing cases.
  6. Lock it in — add the fixed cases to the golden set.

Isolating the stage in RAG

  • Answer not in retrieved chunks → retrieval problem (context recall). Stop editing the generation prompt.
  • Answer in chunks but wrong output → generation or prompt problem.
  • Feeding gold context by hand and seeing if the answer becomes right is the fastest test.

Reading a classifier with precision and recall

For a classification-style LLM task, look per class. The bank's classifier on 1,000 complaints, for the fraud class:

Predicted fraudPredicted other
Actually fraud (100)72 (true positives)28 (false negatives)
Actually not fraud (900)18 (false positives)882 (true negatives)
Text
precision = TP / (TP + FP) = 72 / 90  = 0.80recall    = TP / (TP + FN) = 72 / 100 = 0.72F1        = 2 * P * R / (P + R)       = 0.76accuracy  = (72 + 882) / 1000         = 95.4%

Accuracy looks excellent, but a classifier that never says fraud would already get 90%. Recall 0.72 means 28 fraud cases went to the slow queue. Reading those 28 is the debug step.

Common hidden causes

A truncated context window, a template variable that rendered empty, a chunk boundary splitting a table, a tool returning an error the model ignored, or wrong labels in the eval itself.

A real-life example

The bank reads the 28 missed fraud complaints. Categories: 15 describe an unauthorised UPI debit but never use the word "fraud" (labelled upi); 7 are in Hinglish; 4 are long emails where the key sentence is at the end; 2 are mislabelled in the golden set.

The fixes follow the counts. A labelling-guide rule becomes a prompt rule: "a debit the customer did not authorise is fraud, even if it happened through UPI". Two Hinglish examples are added. Two golden labels are corrected. One change at a time, the full suite re-runs: fraud recall goes 0.72 → 0.84 → 0.88, precision stays at 0.80, and no other class drops more than a point.

Follow-up questions to expect

  • "How do you cluster failures at scale?" — Label a first batch by hand, then have an LLM tag the rest using your categories, or embed failing inputs and cluster; always read samples from each cluster.
  • "What if the model is simply not capable?" — After fixing retrieval and prompts, compare a stronger model on the failing set; if it solves them, it is a capability gap, and you weigh cost.
  • "Precision or recall — which matters more for fraud?" — Usually recall, because a missed fraud case costs more than an extra review; set the threshold from the business cost of each error.