Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Users give contradictory feedback: one user says responses are too short, another says they’re too verbose. How do you convert subjective human feedback into actionable training signals?
What you need to know
From "contradictory" to "conditional"
Feedback without context is two opposite votes. With context it becomes a pattern: developers asking code questions want a short answer and a snippet; first-time users asking about billing want a step-by-step explanation. The preference depends on who is asking and what they asked.
1import pandas as pd23fb = pd.read_parquet("feedback.parquet") # one row per rating, with context columns4table = (fb.groupby(["intent", "segment"])5 .agg(n=("rating", "size"),6 too_long=("reason", lambda r: (r == "too_long").mean()),7 too_short=("reason", lambda r: (r == "too_short").mean()),8 median_words=("answer_words", "median")))9print(table.sort_values("n", ascending=False).head(10))This table shows, for each intent and user segment, how often answers were called too long or too short and how long they actually were. Opposite complaints usually separate cleanly by row.
Better signals
Thumbs and free text
- Noisy, and mostly from unhappy users
- No reference point: too long compared to what?
- Easy to collect
Pairwise comparison
- Two answers to the same prompt, pick one
- Cleaner, relative judgement
- The format preference training uses directly
The pipeline
- Capture context — intent, segment, expertise, device, answer length, and whether the answer was correct.
- Segment — find the conditional patterns.
- Write a rubric — "billing, new user: numbered steps, under 150 words; code, developer: answer plus snippet, under 80 words".
- Enforce cheaply first — routing and system-prompt rules per intent often solve it without training.
- Then train — build preference pairs that follow the rubric, with the context in the prompt so the model learns the condition.
- Offer a control — "Shorter" and "More detail" buttons; each click is a clean labelled signal.
Traps
- Vocal minority. Weight feedback by traffic share, not by who complains most.
- Length bias. In pairwise tests, raters and LLM judges often prefer the longer answer. Keep training on those pairs and everything grows long. Compare answers of similar length or control for length.
- Low agreement. Measure inter-rater agreement (for example Krippendorff's alpha). If humans disagree with each other, the model cannot learn one right answer — make it a product setting instead.
A real-life example
Scenario, numbers made up. A coding-help platform gets 1,800 "too long" and 1,200 "too short" ratings in a month. The team almost adds "be concise" to the system prompt.
Segmenting shows 80% of "too long" comes from experienced developers on debugging questions, while 70% of "too short" comes from students on concept questions ("what is a closure?"). They add an intent router with two answer styles and a "More detail" button. Thumbs-down on length falls by 60%. Clicks on "More detail" become preference data for the next fine-tune, conditioned on user level.
Follow-up questions to expect
- "How do you know the user's expertise?" — Signals like account type, past questions, code in the prompt, or a one-time setting. Keep it simple and let users change it.
- "How many preference pairs do you need?" — It depends on the change; start from a few thousand good pairs for a narrow behaviour, and check for regressions on the gold set.
- "Wouldn't a length control confuse users?" — Keep a good default per intent; the control is for the minority who want something else.