Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

How do you prove a RAG assistant is good enough to launch when there is no eval harness yet?


What you need to know

"Good enough to launch" needs a number and a threshold agreed before you look at results. Otherwise the demo that impressed the VP becomes the launch criterion.

Where the questions come from

Take them from real traffic: support tickets, search logs, emails to the help desk. Questions you invent yourself are cleaner and easier than what users actually type, so a hand-written set flatters the system. Include the awkward ones: typos, two questions in one, questions the documents do not answer.

Score the two halves separately

StageMetricPlain meaningNeeds an LLM?
Retrievalrecall@kWas the evidence chunk in the top k?No
RetrievalMRRHow near the top was it?No
RetrievalContext precisionHow much of what we retrieved was relevant?Usually a judge
GenerationFaithfulnessIs every claim supported by the retrieved text?Judge
GenerationAnswer correctnessDoes it match the reference answer?Judge or human
GenerationAbstention accuracyDoes it say "I don't know" when the documents lack the answer?Rule or judge

Retrieval metrics are cheap and deterministic, so run them on every change. Generation metrics need a judge, which is where care is needed.

Calibrate the judge before trusting it

  1. Label — a person grades 50 answers as correct or incorrect, faithful or not.
  2. Judge — run the LLM judge on the same 50 answers.
  3. Compare — measure agreement; for a less lopsided view, use Cohen's kappa.
  4. Decide — set a bar in advance (for example 85% agreement); below it, fix the judge prompt or keep humans in the loop.

An uncalibrated judge gives you precise-looking numbers with unknown meaning. Libraries such as Ragas provide ready-made metrics, but the calibration step is still yours.

YAML
# ci/rag-eval.yaml — runs on every prompt, chunking or model changedataset: evals/golden_v3.jsonl        # 240 questionsthresholds:  recall_at_5: 0.85  faithfulness: 0.90  answer_correctness: 0.80  abstention_accuracy: 0.90fail_on_regression: 0.02             # block merge if any metric drops by 2 points

Offline is not the whole story

A golden set is frozen; users are not. At launch, track thumbs-down rate, escalation-to-human rate, the share of "I don't know" answers, and follow-up rephrasing (a user asking the same thing twice). Feed every bad production case back into the golden set.

A real-life example

Scenario (illustrative numbers). A telecom company wants to launch a plan-and-billing assistant. The team exports 3,000 recent chat questions, removes duplicates, and samples 240 across 12 intents, including 30 the documents cannot answer. Two support leads write reference answers in three days.

First run: recall@5 is 0.79, faithfulness 0.93, correctness 0.74, abstention 0.55. The model rarely admits it does not know. They add a hybrid search leg and an explicit abstain rule. Second run: recall@5 0.88, correctness 0.83, abstention 0.90. The judge agreed with human labels on 44 of 50 samples (88%), so they trust it for CI. They launch to 10% of users, and the thumbs-down rate matches the offline correctness closely enough to keep the thresholds.

Follow-up questions to expect

  • "How big should the golden set be?" — Big enough to cover every major intent with 15 to 20 examples each; a few hundred is typical. Coverage matters more than size.
  • "Can you generate test questions with an LLM?" — Useful to fill gaps, but keep real questions as the core; synthetic ones tend to be easier and phrased like the documents.
  • "How do you stop the set going stale?" — Add failures from production every week, and re-sample from new traffic every quarter.