Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your eval pipeline catches factual errors but misses tone, politeness, and subtle trust issues. How do you evaluate non-factual qualities in LLM systems reliably?
What you need to know
Decompose before you measure
"Is the tone good?" gets unreliable scores from humans and models alike. Specific dimensions get consistent ones:
| Dimension | A bad example |
|---|---|
| Warmth | "Request processed." to a user reporting a family death |
| Condescension | "As I already explained, it's simple..." |
| Over-apology | Three "I'm so sorry" in a two-line answer |
| Unwarranted certainty | "This will definitely fix it" for a guess |
| Fit to expertise | Explaining what an API is to a senior engineer |
| Appropriate refusal | Refusing a harmless question, or lecturing |
Anchor the scale
For each dimension, write what a 1, 3 and 5 look like, with a real example of each from your own logs. Anchors are what make a judge's scores stable between runs and between raters.
Pairwise with position swap
Absolute 1–5 scores bunch up around 4, so small regressions disappear. Asking "which of these two is better on warmth?" is more sensitive. LLM judges also tend to prefer whichever answer comes first, so ask twice with the order swapped:
1def pairwise(judge, prompt, a, b, dimension):2 first = judge(prompt, a, b, dimension) # returns "A", "B" or "tie"3 second = judge(prompt, b, a, dimension) # same pair, order swapped4 second = {"A": "B", "B": "A"}.get(second, "tie")5 return first if first == second else "tie" # disagreement means no clear winnerA result counts only if it holds in both orders; otherwise it is a tie. This removes position bias from the win rate.
Behavioural proxies from production
Rubrics only catch what you thought of. Production signals catch the rest: users rephrasing, messages like "that's not what I asked", abandonment, and escalation to a human.
Trust is its own measurement
Score two things separately. Calibration: when the answer sounds confident, is it right more often? Group answers by how certain they sound and check accuracy in each group. Citation honesty: do cited sources support the claims? Overconfidence loses trust fastest, and a fact checker only looking at verifiable claims cannot see it.
- Pick dimensions — about six, from real complaints.
- Anchor — examples at 1, 3 and 5 for each.
- Human labels — three or more raters on a few hundred responses.
- Calibrate the judge — keep it only where it tracks human consensus.
- Run pairwise per release — new candidate against production, both orders.
- Watch proxies — rephrase and escalation rates after launch.
A real-life example
Scenario, numbers made up. A lending app's support assistant passes every factual check. But complaints say it sounds "cheerful about my loan default" and "talks down to me".
The team builds a six-dimension rubric with anchors from 60 real conversations. Three support leads label 300 responses; the first judge prompt agrees with their consensus 58% of the time on warmth, and after adding anchors, 84%. Pairwise runs on the next release candidate show it loses on warmth in 31% of collections conversations, so it is held back. After a prompt fix, rephrase rate in collections conversations falls from 14% to 8%.
Follow-up questions to expect
- "What biases do LLM judges have?" — Position bias, a preference for longer answers, and sometimes a preference for text from their own model family. Swap order, control length, and use a different model family as judge where you can.
- "Why at least three human raters?" — One rater's taste is not a standard. With three you can measure agreement and use the majority as consensus.
- "How often do you recalibrate?" — Whenever the judge model or prompt changes, and on a regular sample to catch drift.