Course Content
LLM Evaluation
6 sections · 50 lessons
What are the trade-offs between using strong vs weak models as judges?
What you need to know
What each is good at
Strong judge
- Higher agreement with experts
- Handles multi-criteria rubrics and long context
- Needed for reasoning, code and subtle faithfulness
- Costly; use on samples and offline suites
Small or weak judge
- Cheap enough for high coverage
- Good on narrow checks: PII present? tone polite? cites a source?
- Weak on long transcripts and reasoning
- Can be fine-tuned for one check
Coverage vs accuracy
Catching a live incident depends on both. Suppose a bug makes 5% of answers leak an internal URL. A small judge with 85% recall on 100% of traffic catches most leaking answers within minutes. A perfect judge on 2% of traffic sees only a handful of them per hour. For monitoring, coverage often beats per-item accuracy.
The competence ceiling
A judge verifies by comparing with something. Given a source, even a small model can check whether "15 days" appears. Without a source, a small model asked whether a tax rule is correct gives confident agreement, not verification. Match judge strength to the difficulty of the check, and give judges evidence whenever possible.
The cascade
- Rules — schema, regex, allowed labels, banned phrases. 100% of traffic.
- Small judge — narrow binary checks on high volume. Returns a confidence or can abstain.
- Strong judge — cases the small judge is unsure about, flagged traces, and the offline CI suite.
- Humans — calibration set and the hardest disagreements.
Measure every tier against the same human-labelled set so you know the accuracy you are trading for cost.
A real-life example
The code-review bot's team evaluates 150 human-labelled review comments for "technically correct". A small model agrees with the senior engineers at kappa 0.38; a strong model at kappa 0.74. For "does the comment reference a specific line?", both get above 0.9.
So they split: the small model checks the line-reference rule on every comment; the strong model checks technical correctness on 10% of comments plus every comment a developer marked "not helpful". Monthly judge cost is about a fifth of running the strong model on everything, and the correctness trend still has enough samples to show a regression within a week.
Follow-up questions to expect
- "Can a weaker model judge a stronger one?" — For narrow checks with evidence, yes. For open reasoning quality, it tends to miss errors it cannot solve itself; use a strong judge or humans.
- "Would you use the generator model as the judge?" — Only if necessary; same-family judges show self-preference. Prefer another family, or at least calibrate for this bias.
- "How do you pick the cascade thresholds?" — From the labelled set: choose the small judge's confidence cut-off so that its accepted decisions meet your precision target, and send the rest up.