Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Users rate 2 million answers thumbs-up/down. Now you realize you have no idea how to use that data to improve the model. How do you build a closed-loop system where user feedback genuinely improves the LLM?
What you need to know
Why raw thumbs are hard to use
- Selection bias — a small share of users rate, mostly at the extremes.
- Mixed meanings — thumbs-down can mean wrong, incomplete, slow, off-topic or "I didn't like the policy".
- Inconsistency — the same answer gets different ratings from different people.
- Missing context — a thumb without the query, retrieved chunks and versions cannot be diagnosed.
The loop
- Enrich at capture — store the query, retrieved document IDs, prompt and model versions, latency, and a reason chip (wrong / incomplete / off-topic / tone).
- Cluster the negatives — embed the queries behind thumbs-down and label each cluster by cause.
- Route to the owner — each cause has a cheaper fix than retraining.
- Train only on curated data — preference pairs where a human wrote or approved a better answer.
- Measure the loop — did the thumbs-down rate for that cluster fall after the fix?
| Cause | Owner | Fix |
|---|---|---|
| Right document not retrieved | Search team | Chunking, hybrid search, reranking |
| Answer not in any document | Content team | A gap report: write the missing doc |
| Instruction or format gap | Prompt owner | Prompt change, canaried |
| Genuine model limitation | ML team | Training data |
Training on feedback, carefully
For the model slice, build preference pairs (chosen versus rejected answers) from cases where a reviewer supplied or approved the better answer, and train with DPO on a few thousand clean pairs rather than millions of noisy ones. Methods such as KTO can learn from unpaired good/bad labels, but they still inherit the noise, so filter first. Balance for length so the model does not simply learn to be longer, and evaluate on a held-out set before and after.
1def triage(feedback):2 negatives = [f for f in feedback if f.rating == "down"]3 labels = cluster(embed([f.query for f in negatives]), k=50)4 report = {}5 for f, c in zip(negatives, labels):6 cause = ("retrieval" if not f.gold_doc_retrieved else7 "content_gap" if f.reason == "incomplete" and not f.doc_exists else8 "prompt" if f.reason in ("tone", "off-topic") else "model")9 report.setdefault((c, cause), []).append(f.id)10 return sorted(report.items(), key=lambda kv: -len(kv[1]))In practice the cause labels come from reviewers on a sample of each cluster; the rules above just show how the fields combine.
A real-life example
Scenario, numbers made up. A banking app's assistant has 2 million ratings and 11% thumbs-down. The team adds reason chips and clusters six weeks of negatives into 50 groups. Labelling a sample from each shows roughly 40% retrieval misses (for example, credit-card fee questions landing on debit-card documents), 30% questions no document answers, 20% prompt and format issues, and 10% real model errors.
In one month they fix metadata filters, publish 35 missing help articles, and adjust the prompt for fee tables. Thumbs-down falls from 11% to 7% before any training. Then they build 4,000 reviewed preference pairs for the remaining model errors and run DPO; the held-out eval improves, and the length balance keeps answers the same size.
Follow-up questions to expect
- "Why not train on thumbs-up answers directly?" — A thumbs-up often means "fast and polite", not "correct". Training on it can teach style over accuracy.
- "How do you get more useful feedback?" — Ask for a reason with one tap, and use implicit signals such as rephrasing, copying the answer or escalating to a human.
- "How do you prove the loop works?" — Track the thumbs-down rate per cluster before and after each fix, and keep a held-out eval for model changes.