Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your automated eval scores 92%, but customers complain daily. The eval set is a year old and doesn't reflect real usage. How do you build an eval set that actually mirrors production?
What you need to know
Why an old eval set lies
An eval set is a sample of the questions you expect. Users change over a year: new features launch, new customer types arrive, and people learn what the bot is good at and ask harder things. This is distribution shift, and it is why a score can stay at 92% while quality falls.
There is a second trap. A random sample of logs is dominated by easy, frequent "head" queries ("where is my order?"). The hard "tail" questions, where complaints come from, appear only a few times each. A random set scores well because it is mostly easy.
Build it from production, stratified
- Pull real traffic — take 30 days of logs, redact PII before anything is stored for eval.
- Cluster by meaning — embed each query and group them into topics, so every kind of question is visible.
- Sample every cluster — take a fixed number per cluster, so rare topics are covered.
- Over-sample failures — weight toward thumbs-down, escalation to a human, abandonment, and rephrased retries.
- Add hard cases — ambiguous, out-of-scope, prompt-injection and "the documents cannot answer this" questions.
- Label and version — human reference answers for a core set, a calibrated LLM judge for the rest, and a version number on the set.
A reasonable mix:
| Slice | Share | Why it is there |
|---|---|---|
| Representative traffic | ~60% | Tells you what the average user feels |
| Known failure modes from complaints | ~25% | Stops old bugs from coming back |
| Adversarial and edge cases | ~15% | Tests refusal, safety and "I don't know" |
A sketch of the sampling step:
1import numpy as np2from sklearn.cluster import KMeans34def sample_eval_cases(queries, embeddings, failed, n_clusters=40, per_cluster=15, boost=3.0):5 """failed: bool array — thumbs-down, escalated, or rephrased within 2 minutes."""6 labels = KMeans(n_clusters=n_clusters, n_init="auto", random_state=0).fit_predict(embeddings)7 rng = np.random.default_rng(0)8 picked = []9 for c in range(n_clusters):10 idx = np.where(labels == c)[0]11 w = np.where(failed[idx], boost, 1.0)12 k = min(per_cluster, len(idx))13 picked += rng.choice(idx, size=k, replace=False, p=w / w.sum()).tolist()14 return [(queries[i], int(labels[i])) for i in picked]Taking the same number per cluster guarantees the tail is covered. Because that over-weights rare topics, report two numbers: a traffic-weighted score (weight each cluster by its real share of queries) and the worst cluster's score, which shows where users are suffering.
Labelling at scale
Humans write reference answers for a core of about 200 cases. For the rest, an LLM judge (a model that grades answers against a rubric) scales cheaply, but only after you check it: have humans grade 50–100 answers and measure how often the judge agrees. If agreement is below roughly 85%, fix the rubric before trusting the judge.
Keep it alive
The old set failed because it was built once. Make it a process: a weekly job pushes new low-rated conversations into a review queue, every confirmed complaint becomes a permanent test case, and the set is versioned so you only compare scores on the same version. Track whether the offline score moves with your online quality metric (complaint rate, thumbs-down rate). If offline goes up while complaints also go up, the eval set is wrong again.
A real-life example
Scenario, numbers made up. A food-delivery app's support bot scores 92% on an 800-question eval set written at launch. Most of those questions are about order status. Since then, the app has added scheduled orders, UPI Autopay for a membership plan, and a Hindi-English chat option.
The team embeds 30 days of conversations (about 1.2 million) into 40 clusters. Six clusters have no eval coverage at all — refunds to UPI, membership cancellation, and Hinglish questions among them — and those six produce 45% of the thumbs-down. They build a 1,000-case set with the mix above. The bot scores 71% on it. That drop is good news: the eval now sees what customers see.
Two sprints of fixes follow (better retrieval for membership docs, a Hinglish test slice, a refusal rule for policy questions the docs do not cover). The score reaches 84%, and complaint tickets fall by about a third. The weekly job adds roughly 30 new cases a week.
Follow-up questions to expect
- "How big should the eval set be?" — Big enough that the change you care about is larger than the noise. With 1,000 cases, a 2-point move is roughly at the edge of noise; for smaller effects, add cases or compare paired results on the same questions.
- "How do you stop the team overfitting to the eval set?" — Keep a held-out split that nobody tunes prompts against, and refresh the main set with new production cases every month.
- "Is it safe to put production data in an eval set?" — Only after PII redaction and with the retention rules your privacy policy allows; store the redacted text, not the raw logs.