LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

What are the trade-offs between using strong vs weak models as judges?


The judge cascade for review commentsRules on every commentSmall judge: cites a line? kappa above 0.9Strong judge: technically correct? kappa 0.74Engineers: calibration set and disputes
Each tier handles only what the cheaper tier below cannot, so strong-judge spend goes where the small one scored kappa 0.38.

What you need to know

What each is good at

Strong judge

  • Higher agreement with experts
  • Handles multi-criteria rubrics and long context
  • Needed for reasoning, code and subtle faithfulness
  • Costly; use on samples and offline suites

Small or weak judge

  • Cheap enough for high coverage
  • Good on narrow checks: PII present? tone polite? cites a source?
  • Weak on long transcripts and reasoning
  • Can be fine-tuned for one check

Coverage vs accuracy

Catching a live incident depends on both. Suppose a bug makes 5% of answers leak an internal URL. A small judge with 85% recall on 100% of traffic catches most leaking answers within minutes. A perfect judge on 2% of traffic sees only a handful of them per hour. For monitoring, coverage often beats per-item accuracy.

The competence ceiling

A judge verifies by comparing with something. Given a source, even a small model can check whether "15 days" appears. Without a source, a small model asked whether a tax rule is correct gives confident agreement, not verification. Match judge strength to the difficulty of the check, and give judges evidence whenever possible.

The cascade

  1. Rules — schema, regex, allowed labels, banned phrases. 100% of traffic.
  2. Small judge — narrow binary checks on high volume. Returns a confidence or can abstain.
  3. Strong judge — cases the small judge is unsure about, flagged traces, and the offline CI suite.
  4. Humans — calibration set and the hardest disagreements.

Measure every tier against the same human-labelled set so you know the accuracy you are trading for cost.

A real-life example

The code-review bot's team evaluates 150 human-labelled review comments for "technically correct". A small model agrees with the senior engineers at kappa 0.38; a strong model at kappa 0.74. For "does the comment reference a specific line?", both get above 0.9.

So they split: the small model checks the line-reference rule on every comment; the strong model checks technical correctness on 10% of comments plus every comment a developer marked "not helpful". Monthly judge cost is about a fifth of running the strong model on everything, and the correctness trend still has enough samples to show a regression within a week.

Follow-up questions to expect

  • "Can a weaker model judge a stronger one?" — For narrow checks with evidence, yes. For open reasoning quality, it tends to miss errors it cannot solve itself; use a strong judge or humans.
  • "Would you use the generator model as the judge?" — Only if necessary; same-family judges show self-preference. Prefer another family, or at least calibrate for this bias.
  • "How do you pick the cascade thresholds?" — From the labelled set: choose the small judge's confidence cut-off so that its accepted decisions meet your precision target, and send the rest up.