Course Content
LLM Evaluation
6 sections · 50 lessons
How do you evaluate the robustness of a model across input variations?
What you need to know
Perturbation types
| Type | Examples |
|---|---|
| Surface | Typos, ALL CAPS, missing punctuation, extra spaces, emoji |
| Paraphrase | Reworded question, formal versus casual |
| Language | Hindi, Hinglish, regional English, translations |
| Structure | Key fact at the start versus end, added irrelevant sentences |
| Format | Prose versus bullet list versus a forwarded email with signatures |
| Adversarial | Plausible but misleading context, conflicting instructions |
Two numbers
- Perturbed accuracy — the score on all variants.
- Flip rate — the share of items whose result changes between the original and a variant.
A system with a stable 80% is more trustworthy than one averaging 85% whose answer flips on a third of paraphrases: users cannot predict it.
Keep noise out
Run the unperturbed original several times first. If it already flips 5% of the time on its own, only flips above that level are caused by the perturbation.
Generating variants
Rules for surface changes (typo injection, casing), an LLM for paraphrases and translations, and a native speaker to check a sample, because generated Hinglish is often unnatural.
A real-life example
The bank takes 300 complaints from its golden set and creates five variants each: typos, all caps, Hinglish, a forwarded-email wrapper with signatures and disclaimers, and the key sentence moved to the end.
| Variant | Accuracy | Flip rate vs original |
|---|---|---|
| Original | 88% | 3% (run-to-run noise) |
| Typos | 87% | 4% |
| All caps | 86% | 5% |
| Hinglish | 79% | 14% |
| Forwarded email | 81% | 11% |
| Key sentence at end | 84% | 8% |
Typos and caps are within noise. Hinglish and forwarded emails are real weaknesses. The team adds Hinglish examples and a pre-processing step that strips signatures and disclaimers. Flip rates for those variants drop to 6% and 5%, and the variants stay in the regression suite.
Follow-up questions to expect
- "How is robustness different from adversarial testing?" — Robustness uses natural variations real users produce; adversarial testing uses deliberate attacks. Both use perturbations, with different intent.
- "How do you know a paraphrase kept the meaning?" — Check a sample by hand, or use an entailment check in both directions; discard variants that changed the meaning.
- "What about long-context position effects?" — Move the key passage to the start, middle and end of the context; many models are weaker when it is in the middle.