Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your automated eval scores 92%, but customers complain daily. The eval set is a year old and doesn't reflect real usage. How do you build an eval set that actually mirrors production?


Rebuilding the eval set from real traffic30 days oflogs, PIIredactedClusterqueriesby meaningSample everycluster,boost failuresAdd adversarial andunanswerable casesLabel, version,refresh weeklySix clusters had no coverage and caused 45 percent of thumbs-down.
The score fell from 92 to 71 percent on the new set, and that drop was the first honest number the team had seen all year.

What you need to know

Why an old eval set lies

An eval set is a sample of the questions you expect. Users change over a year: new features launch, new customer types arrive, and people learn what the bot is good at and ask harder things. This is distribution shift, and it is why a score can stay at 92% while quality falls.

There is a second trap. A random sample of logs is dominated by easy, frequent "head" queries ("where is my order?"). The hard "tail" questions, where complaints come from, appear only a few times each. A random set scores well because it is mostly easy.

Build it from production, stratified

  1. Pull real traffic — take 30 days of logs, redact PII before anything is stored for eval.
  2. Cluster by meaning — embed each query and group them into topics, so every kind of question is visible.
  3. Sample every cluster — take a fixed number per cluster, so rare topics are covered.
  4. Over-sample failures — weight toward thumbs-down, escalation to a human, abandonment, and rephrased retries.
  5. Add hard cases — ambiguous, out-of-scope, prompt-injection and "the documents cannot answer this" questions.
  6. Label and version — human reference answers for a core set, a calibrated LLM judge for the rest, and a version number on the set.

A reasonable mix:

SliceShareWhy it is there
Representative traffic~60%Tells you what the average user feels
Known failure modes from complaints~25%Stops old bugs from coming back
Adversarial and edge cases~15%Tests refusal, safety and "I don't know"

A sketch of the sampling step:

Python
import numpy as npfrom sklearn.cluster import KMeansdef sample_eval_cases(queries, embeddings, failed, n_clusters=40, per_cluster=15, boost=3.0):    """failed: bool array — thumbs-down, escalated, or rephrased within 2 minutes."""    labels = KMeans(n_clusters=n_clusters, n_init="auto", random_state=0).fit_predict(embeddings)    rng = np.random.default_rng(0)    picked = []    for c in range(n_clusters):        idx = np.where(labels == c)[0]        w = np.where(failed[idx], boost, 1.0)        k = min(per_cluster, len(idx))        picked += rng.choice(idx, size=k, replace=False, p=w / w.sum()).tolist()    return [(queries[i], int(labels[i])) for i in picked]

Taking the same number per cluster guarantees the tail is covered. Because that over-weights rare topics, report two numbers: a traffic-weighted score (weight each cluster by its real share of queries) and the worst cluster's score, which shows where users are suffering.

Labelling at scale

Humans write reference answers for a core of about 200 cases. For the rest, an LLM judge (a model that grades answers against a rubric) scales cheaply, but only after you check it: have humans grade 50–100 answers and measure how often the judge agrees. If agreement is below roughly 85%, fix the rubric before trusting the judge.

Keep it alive

The old set failed because it was built once. Make it a process: a weekly job pushes new low-rated conversations into a review queue, every confirmed complaint becomes a permanent test case, and the set is versioned so you only compare scores on the same version. Track whether the offline score moves with your online quality metric (complaint rate, thumbs-down rate). If offline goes up while complaints also go up, the eval set is wrong again.

A real-life example

Scenario, numbers made up. A food-delivery app's support bot scores 92% on an 800-question eval set written at launch. Most of those questions are about order status. Since then, the app has added scheduled orders, UPI Autopay for a membership plan, and a Hindi-English chat option.

The team embeds 30 days of conversations (about 1.2 million) into 40 clusters. Six clusters have no eval coverage at all — refunds to UPI, membership cancellation, and Hinglish questions among them — and those six produce 45% of the thumbs-down. They build a 1,000-case set with the mix above. The bot scores 71% on it. That drop is good news: the eval now sees what customers see.

Two sprints of fixes follow (better retrieval for membership docs, a Hinglish test slice, a refusal rule for policy questions the docs do not cover). The score reaches 84%, and complaint tickets fall by about a third. The weekly job adds roughly 30 new cases a week.

Follow-up questions to expect

  • "How big should the eval set be?" — Big enough that the change you care about is larger than the noise. With 1,000 cases, a 2-point move is roughly at the edge of noise; for smaller effects, add cases or compare paired results on the same questions.
  • "How do you stop the team overfitting to the eval set?" — Keep a held-out split that nobody tunes prompts against, and refresh the main set with new production cases every month.
  • "Is it safe to put production data in an eval set?" — Only after PII redaction and with the retention rules your privacy policy allows; store the redacted text, not the raw logs.