Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your benchmark scores improved after fine-tuning, but real user satisfaction dropped sharply. How do you design evaluations that correlate with actual human experience?
What you need to know
Why the benchmark and users disagree
A public benchmark is clean, short and scored on correctness. Your users send messy, ambiguous questions and care about how the answer feels. Fine-tuning often shifts behaviour the benchmark cannot see: answers get longer, the model refuses harmless requests more, or it stops asking clarifying questions. The score goes up; the experience goes down.
Offline and online measures
| Layer | Examples | Role |
|---|---|---|
| Offline gold set | A few hundred real queries, labelled by people like your users | Gates a release |
| Offline extra axes | Length, refusal rate on harmless requests, tone, past-complaint regression suite | Catches what the benchmark ignores |
| Online outcomes | Task completion, rephrase rate, escalation to a human, thumbs, return rate | Decides a release |
Validate the eval itself
An LLM judge or automated score is only useful if it moves with human judgement. Check that directly:
1from scipy.stats import spearmanr23# one row per sampled conversation: automated score and a human satisfaction label (1-5)4judge = [r.judge_score for r in sample]5human = [r.human_rating for r in sample]6rho, p = spearmanr(judge, human)7print(f"rank correlation {rho:.2f} (p={p:.3f}) on {len(sample)} conversations")Spearman rank correlation asks: when humans rate A above B, does the judge too? A common rule of thumb is not to rely on a judge whose correlation with humans is below about 0.7, but set your own bar from the decisions you need to make.
The rebuild
- Sample production traffic — stratified by intent, user segment and difficulty, deliberately including the messy tail.
- Label with real-user proxies — people who resemble your users, not only engineers.
- Add the missing axes — length, harmless-request refusals, tone, and a regression suite built from past complaints.
- Check correlation — every automated metric against a human outcome.
- Ship behind a flag — offline evals gate; an A/B at 5% with guardrail metrics like rephrase rate and abandonment decides.
- Refresh monthly — add new traffic, keep a held-out slice you never tune against.
A real-life example
Scenario, numbers made up. An edtech company fine-tunes its doubt-solving model. Its benchmark score rises 6 points. Two weeks after launch, app ratings drop and the share of students who ask the same question again rises from 11% to 19%.
The team samples 400 real doubts and has teachers rate old and new answers. The new model's answers are twice as long, it refuses 4% of harmless questions about exam strategy, and it no longer asks "which chapter?" when a question is vague. Their LLM judge correlates only 0.35 with teacher ratings. They rebuild the judge rubric around what teachers valued, add length and refusal checks, and retrain; the next version gains 2 benchmark points instead of 6, but the repeat-question rate falls to 9%.
Follow-up questions to expect
- "How big should the gold set be?" — Big enough per slice to see the differences you care about: often a few hundred items overall, with at least dozens per important slice.
- "How do you avoid overfitting to the gold set?" — Keep a held-out slice nobody tunes against, and refresh from new traffic regularly.
- "What if you can't run an A/B test?" — Use a shadow comparison: run both models on live traffic, show one, and have humans rate a sample of pairs.