Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your eval pipeline catches factual errors but misses tone, politeness, and subtle trust issues. How do you evaluate non-factual qualities in LLM systems reliably?


"Tone" broken into things a judge can scoreToneWarmthCondescensionOver-apologyUnwarranted certaintyFit to user expertiseAppropriate refusal
Nobody agrees on a single tone score, but raters and judges converge once each dimension has anchored examples.

What you need to know

Decompose before you measure

"Is the tone good?" gets unreliable scores from humans and models alike. Specific dimensions get consistent ones:

DimensionA bad example
Warmth"Request processed." to a user reporting a family death
Condescension"As I already explained, it's simple..."
Over-apologyThree "I'm so sorry" in a two-line answer
Unwarranted certainty"This will definitely fix it" for a guess
Fit to expertiseExplaining what an API is to a senior engineer
Appropriate refusalRefusing a harmless question, or lecturing

Anchor the scale

For each dimension, write what a 1, 3 and 5 look like, with a real example of each from your own logs. Anchors are what make a judge's scores stable between runs and between raters.

Pairwise with position swap

Absolute 1–5 scores bunch up around 4, so small regressions disappear. Asking "which of these two is better on warmth?" is more sensitive. LLM judges also tend to prefer whichever answer comes first, so ask twice with the order swapped:

Python
def pairwise(judge, prompt, a, b, dimension):    first = judge(prompt, a, b, dimension)    # returns "A", "B" or "tie"    second = judge(prompt, b, a, dimension)   # same pair, order swapped    second = {"A": "B", "B": "A"}.get(second, "tie")    return first if first == second else "tie"   # disagreement means no clear winner

A result counts only if it holds in both orders; otherwise it is a tie. This removes position bias from the win rate.

Behavioural proxies from production

Rubrics only catch what you thought of. Production signals catch the rest: users rephrasing, messages like "that's not what I asked", abandonment, and escalation to a human.

Trust is its own measurement

Score two things separately. Calibration: when the answer sounds confident, is it right more often? Group answers by how certain they sound and check accuracy in each group. Citation honesty: do cited sources support the claims? Overconfidence loses trust fastest, and a fact checker only looking at verifiable claims cannot see it.

  1. Pick dimensions — about six, from real complaints.
  2. Anchor — examples at 1, 3 and 5 for each.
  3. Human labels — three or more raters on a few hundred responses.
  4. Calibrate the judge — keep it only where it tracks human consensus.
  5. Run pairwise per release — new candidate against production, both orders.
  6. Watch proxies — rephrase and escalation rates after launch.

A real-life example

Scenario, numbers made up. A lending app's support assistant passes every factual check. But complaints say it sounds "cheerful about my loan default" and "talks down to me".

The team builds a six-dimension rubric with anchors from 60 real conversations. Three support leads label 300 responses; the first judge prompt agrees with their consensus 58% of the time on warmth, and after adding anchors, 84%. Pairwise runs on the next release candidate show it loses on warmth in 31% of collections conversations, so it is held back. After a prompt fix, rephrase rate in collections conversations falls from 14% to 8%.

Follow-up questions to expect

  • "What biases do LLM judges have?" — Position bias, a preference for longer answers, and sometimes a preference for text from their own model family. Swap order, control length, and use a different model family as judge where you can.
  • "Why at least three human raters?" — One rater's taste is not a standard. With three you can measure agreement and use the majority as consensus.
  • "How often do you recalibrate?" — Whenever the judge model or prompt changes, and on a regular sample to catch drift.