Course Content
LLM Evaluation
6 sections · 50 lessons
How do you debug poor model performance systematically?
What you need to know
Error analysis
- Collect failures — 20 to 50 from the eval or from production thumbs-down, not one anecdote.
- Read full traces — the final prompt after templating, retrieved chunks, tool calls and outputs.
- Label and count — free-text notes first, then categories.
- Isolate the stage — retrieval, prompt, model, tool, post-processing.
- Change one thing — and re-run the whole suite, not only the failing cases.
- Lock it in — add the fixed cases to the golden set.
Isolating the stage in RAG
- Answer not in retrieved chunks → retrieval problem (context recall). Stop editing the generation prompt.
- Answer in chunks but wrong output → generation or prompt problem.
- Feeding gold context by hand and seeing if the answer becomes right is the fastest test.
Reading a classifier with precision and recall
For a classification-style LLM task, look per class. The bank's classifier on 1,000 complaints, for the fraud class:
| Predicted fraud | Predicted other | |
|---|---|---|
| Actually fraud (100) | 72 (true positives) | 28 (false negatives) |
| Actually not fraud (900) | 18 (false positives) | 882 (true negatives) |
precision = TP / (TP + FP) = 72 / 90 = 0.80recall = TP / (TP + FN) = 72 / 100 = 0.72F1 = 2 * P * R / (P + R) = 0.76accuracy = (72 + 882) / 1000 = 95.4%Accuracy looks excellent, but a classifier that never says fraud would already get 90%. Recall 0.72 means 28 fraud cases went to the slow queue. Reading those 28 is the debug step.
Common hidden causes
A truncated context window, a template variable that rendered empty, a chunk boundary splitting a table, a tool returning an error the model ignored, or wrong labels in the eval itself.
A real-life example
The bank reads the 28 missed fraud complaints. Categories: 15 describe an unauthorised UPI debit but never use the word "fraud" (labelled upi); 7 are in Hinglish; 4 are long emails where the key sentence is at the end; 2 are mislabelled in the golden set.
The fixes follow the counts. A labelling-guide rule becomes a prompt rule: "a debit the customer did not authorise is fraud, even if it happened through UPI". Two Hinglish examples are added. Two golden labels are corrected. One change at a time, the full suite re-runs: fraud recall goes 0.72 → 0.84 → 0.88, precision stays at 0.80, and no other class drops more than a point.
Follow-up questions to expect
- "How do you cluster failures at scale?" — Label a first batch by hand, then have an LLM tag the rest using your categories, or embed failing inputs and cluster; always read samples from each cluster.
- "What if the model is simply not capable?" — After fixing retrieval and prompts, compare a stronger model on the failing set; if it solves them, it is a capability gap, and you weigh cost.
- "Precision or recall — which matters more for fraud?" — Usually recall, because a missed fraud case costs more than an extra review; set the threshold from the business cost of each error.