Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Users rate 2 million answers thumbs-up/down. Now you realize you have no idea how to use that data to improve the model. How do you build a closed-loop system where user feedback genuinely improves the LLM?


Routing a thumbs-down cluster to its ownerThumbs-down clusterwith reason chipsRetrieval miss: search teamNo document: content teamFormat or tone: prompt ownerModel limit:curated DPO pairs
Most complaints were fixed before any training, because only one slice in ten was a problem the model itself could learn away.

What you need to know

Why raw thumbs are hard to use

  • Selection bias — a small share of users rate, mostly at the extremes.
  • Mixed meanings — thumbs-down can mean wrong, incomplete, slow, off-topic or "I didn't like the policy".
  • Inconsistency — the same answer gets different ratings from different people.
  • Missing context — a thumb without the query, retrieved chunks and versions cannot be diagnosed.

The loop

  1. Enrich at capture — store the query, retrieved document IDs, prompt and model versions, latency, and a reason chip (wrong / incomplete / off-topic / tone).
  2. Cluster the negatives — embed the queries behind thumbs-down and label each cluster by cause.
  3. Route to the owner — each cause has a cheaper fix than retraining.
  4. Train only on curated data — preference pairs where a human wrote or approved a better answer.
  5. Measure the loop — did the thumbs-down rate for that cluster fall after the fix?
CauseOwnerFix
Right document not retrievedSearch teamChunking, hybrid search, reranking
Answer not in any documentContent teamA gap report: write the missing doc
Instruction or format gapPrompt ownerPrompt change, canaried
Genuine model limitationML teamTraining data

Training on feedback, carefully

For the model slice, build preference pairs (chosen versus rejected answers) from cases where a reviewer supplied or approved the better answer, and train with DPO on a few thousand clean pairs rather than millions of noisy ones. Methods such as KTO can learn from unpaired good/bad labels, but they still inherit the noise, so filter first. Balance for length so the model does not simply learn to be longer, and evaluate on a held-out set before and after.

Python
def triage(feedback):    negatives = [f for f in feedback if f.rating == "down"]    labels = cluster(embed([f.query for f in negatives]), k=50)    report = {}    for f, c in zip(negatives, labels):        cause = ("retrieval" if not f.gold_doc_retrieved else                 "content_gap" if f.reason == "incomplete" and not f.doc_exists else                 "prompt" if f.reason in ("tone", "off-topic") else "model")        report.setdefault((c, cause), []).append(f.id)    return sorted(report.items(), key=lambda kv: -len(kv[1]))

In practice the cause labels come from reviewers on a sample of each cluster; the rules above just show how the fields combine.

A real-life example

Scenario, numbers made up. A banking app's assistant has 2 million ratings and 11% thumbs-down. The team adds reason chips and clusters six weeks of negatives into 50 groups. Labelling a sample from each shows roughly 40% retrieval misses (for example, credit-card fee questions landing on debit-card documents), 30% questions no document answers, 20% prompt and format issues, and 10% real model errors.

In one month they fix metadata filters, publish 35 missing help articles, and adjust the prompt for fee tables. Thumbs-down falls from 11% to 7% before any training. Then they build 4,000 reviewed preference pairs for the remaining model errors and run DPO; the held-out eval improves, and the length balance keeps answers the same size.

Follow-up questions to expect

  • "Why not train on thumbs-up answers directly?" — A thumbs-up often means "fast and polite", not "correct". Training on it can teach style over accuracy.
  • "How do you get more useful feedback?" — Ask for a reason with one tap, and use implicit signals such as rephrasing, copying the answer or escalating to a human.
  • "How do you prove the loop works?" — Track the thumbs-down rate per cluster before and after each fix, and keep a held-out eval for model changes.