Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Users give contradictory feedback: one user says responses are too short, another says they’re too verbose. How do you convert subjective human feedback into actionable training signals?


The same complaint split by who asked whattoo longabout rightabout righttoo shortDebugging questionConcept questionExperienced developerStudent
Opposite complaints stop contradicting each other once feedback carries its context, and each cell gets its own rule.

What you need to know

From "contradictory" to "conditional"

Feedback without context is two opposite votes. With context it becomes a pattern: developers asking code questions want a short answer and a snippet; first-time users asking about billing want a step-by-step explanation. The preference depends on who is asking and what they asked.

Python
import pandas as pdfb = pd.read_parquet("feedback.parquet")   # one row per rating, with context columnstable = (fb.groupby(["intent", "segment"])           .agg(n=("rating", "size"),                too_long=("reason", lambda r: (r == "too_long").mean()),                too_short=("reason", lambda r: (r == "too_short").mean()),                median_words=("answer_words", "median")))print(table.sort_values("n", ascending=False).head(10))

This table shows, for each intent and user segment, how often answers were called too long or too short and how long they actually were. Opposite complaints usually separate cleanly by row.

Better signals

Thumbs and free text

  • Noisy, and mostly from unhappy users
  • No reference point: too long compared to what?
  • Easy to collect

Pairwise comparison

  • Two answers to the same prompt, pick one
  • Cleaner, relative judgement
  • The format preference training uses directly

The pipeline

  1. Capture context — intent, segment, expertise, device, answer length, and whether the answer was correct.
  2. Segment — find the conditional patterns.
  3. Write a rubric — "billing, new user: numbered steps, under 150 words; code, developer: answer plus snippet, under 80 words".
  4. Enforce cheaply first — routing and system-prompt rules per intent often solve it without training.
  5. Then train — build preference pairs that follow the rubric, with the context in the prompt so the model learns the condition.
  6. Offer a control — "Shorter" and "More detail" buttons; each click is a clean labelled signal.

Traps

  • Vocal minority. Weight feedback by traffic share, not by who complains most.
  • Length bias. In pairwise tests, raters and LLM judges often prefer the longer answer. Keep training on those pairs and everything grows long. Compare answers of similar length or control for length.
  • Low agreement. Measure inter-rater agreement (for example Krippendorff's alpha). If humans disagree with each other, the model cannot learn one right answer — make it a product setting instead.

A real-life example

Scenario, numbers made up. A coding-help platform gets 1,800 "too long" and 1,200 "too short" ratings in a month. The team almost adds "be concise" to the system prompt.

Segmenting shows 80% of "too long" comes from experienced developers on debugging questions, while 70% of "too short" comes from students on concept questions ("what is a closure?"). They add an intent router with two answer styles and a "More detail" button. Thumbs-down on length falls by 60%. Clicks on "More detail" become preference data for the next fine-tune, conditioned on user level.

Follow-up questions to expect

  • "How do you know the user's expertise?" — Signals like account type, past questions, code in the prompt, or a one-time setting. Keep it simple and let users change it.
  • "How many preference pairs do you need?" — It depends on the change; start from a few thousand good pairs for a narrow behaviour, and check for regressions on the gold set.
  • "Wouldn't a length control confuse users?" — Keep a good default per intent; the control is for the minority who want something else.