Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your benchmark scores improved after fine-tuning, but real user satisfaction dropped sharply. How do you design evaluations that correlate with actual human experience?


What you need to know

Why the benchmark and users disagree

A public benchmark is clean, short and scored on correctness. Your users send messy, ambiguous questions and care about how the answer feels. Fine-tuning often shifts behaviour the benchmark cannot see: answers get longer, the model refuses harmless requests more, or it stops asking clarifying questions. The score goes up; the experience goes down.

Offline and online measures

LayerExamplesRole
Offline gold setA few hundred real queries, labelled by people like your usersGates a release
Offline extra axesLength, refusal rate on harmless requests, tone, past-complaint regression suiteCatches what the benchmark ignores
Online outcomesTask completion, rephrase rate, escalation to a human, thumbs, return rateDecides a release

Validate the eval itself

An LLM judge or automated score is only useful if it moves with human judgement. Check that directly:

Python
from scipy.stats import spearmanr# one row per sampled conversation: automated score and a human satisfaction label (1-5)judge = [r.judge_score for r in sample]human = [r.human_rating for r in sample]rho, p = spearmanr(judge, human)print(f"rank correlation {rho:.2f} (p={p:.3f}) on {len(sample)} conversations")

Spearman rank correlation asks: when humans rate A above B, does the judge too? A common rule of thumb is not to rely on a judge whose correlation with humans is below about 0.7, but set your own bar from the decisions you need to make.

The rebuild

  1. Sample production traffic — stratified by intent, user segment and difficulty, deliberately including the messy tail.
  2. Label with real-user proxies — people who resemble your users, not only engineers.
  3. Add the missing axes — length, harmless-request refusals, tone, and a regression suite built from past complaints.
  4. Check correlation — every automated metric against a human outcome.
  5. Ship behind a flag — offline evals gate; an A/B at 5% with guardrail metrics like rephrase rate and abandonment decides.
  6. Refresh monthly — add new traffic, keep a held-out slice you never tune against.

A real-life example

Scenario, numbers made up. An edtech company fine-tunes its doubt-solving model. Its benchmark score rises 6 points. Two weeks after launch, app ratings drop and the share of students who ask the same question again rises from 11% to 19%.

The team samples 400 real doubts and has teachers rate old and new answers. The new model's answers are twice as long, it refuses 4% of harmless questions about exam strategy, and it no longer asks "which chapter?" when a question is vague. Their LLM judge correlates only 0.35 with teacher ratings. They rebuild the judge rubric around what teachers valued, add length and refusal checks, and retrain; the next version gains 2 benchmark points instead of 6, but the repeat-question rate falls to 9%.

Follow-up questions to expect

  • "How big should the gold set be?" — Big enough per slice to see the differences you care about: often a few hundred items overall, with at least dozens per important slice.
  • "How do you avoid overfitting to the gold set?" — Keep a held-out slice nobody tunes against, and refresh from new traffic regularly.
  • "What if you can't run an A/B test?" — Use a shadow comparison: run both models on live traffic, show one, and have humans rate a sample of pairs.