LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

How do you evaluate the robustness of a model across input variations?


Flip rate against the original, by variant3%4%5%14%11%8%012345originalrerun: noiseHinglishforwarded emailVariants in order: rerun, typos, all caps, Hinglish, forwarded email, key sentence at end.
Typos and caps sit inside run-to-run noise; only the variants well above 3% are real weaknesses.

What you need to know

Perturbation types

TypeExamples
SurfaceTypos, ALL CAPS, missing punctuation, extra spaces, emoji
ParaphraseReworded question, formal versus casual
LanguageHindi, Hinglish, regional English, translations
StructureKey fact at the start versus end, added irrelevant sentences
FormatProse versus bullet list versus a forwarded email with signatures
AdversarialPlausible but misleading context, conflicting instructions

Two numbers

  • Perturbed accuracy — the score on all variants.
  • Flip rate — the share of items whose result changes between the original and a variant.

A system with a stable 80% is more trustworthy than one averaging 85% whose answer flips on a third of paraphrases: users cannot predict it.

Keep noise out

Run the unperturbed original several times first. If it already flips 5% of the time on its own, only flips above that level are caused by the perturbation.

Generating variants

Rules for surface changes (typo injection, casing), an LLM for paraphrases and translations, and a native speaker to check a sample, because generated Hinglish is often unnatural.

A real-life example

The bank takes 300 complaints from its golden set and creates five variants each: typos, all caps, Hinglish, a forwarded-email wrapper with signatures and disclaimers, and the key sentence moved to the end.

VariantAccuracyFlip rate vs original
Original88%3% (run-to-run noise)
Typos87%4%
All caps86%5%
Hinglish79%14%
Forwarded email81%11%
Key sentence at end84%8%

Typos and caps are within noise. Hinglish and forwarded emails are real weaknesses. The team adds Hinglish examples and a pre-processing step that strips signatures and disclaimers. Flip rates for those variants drop to 6% and 5%, and the variants stay in the regression suite.

Follow-up questions to expect

  • "How is robustness different from adversarial testing?" — Robustness uses natural variations real users produce; adversarial testing uses deliberate attacks. Both use perturbations, with different intent.
  • "How do you know a paraphrase kept the meaning?" — Check a sample by hand, or use an entailment check in both directions; discard variants that changed the meaning.
  • "What about long-context position effects?" — Move the key passage to the start, middle and end of the context; many models are weaker when it is in the middle.