Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your reward model prefers 'helpful' answers. The LLM learns to sound confident and verbose — but becomes factually worse. How do you design reward signals that don't get gamed?


What the policy was paid forOne helpfulness score• Long, confident answers win pairs• Length grew from 180 to 420 tokens• Lawyer-graded accuracy fell to 74%• Reward kept rising the whole timeGated, decomposed reward• Unsupported claim means zero reward• Length-balanced preference pairs• KL penalty and early stop on humans• 86% accuracy at about 210 tokens
A rising proxy reward next to a falling human eval is the signature of reward hacking, and the crossing point is the stop condition.

What you need to know

How a reward gets gamed

In RLHF (reinforcement learning from human feedback), people compare pairs of answers, a reward model learns to predict which one they prefer, and the LLM is trained to score high on that reward model. The reward model is a proxy for "good". If long, confident answers were usually preferred, the reward model learns "long and confident is good", and the LLM happily produces that — even when it is wrong.

Length bias is one of the most widely reported cases: preference models trained without controls tend to reward longer answers.

Four defences

DefenceWhat it doesTrade-off
Decomposed reward with a gateSeparate scores; wrong answers get zeroMore labelling and more models to maintain
Length controlBalance pairs so the chosen answer is not usually longer, or penalise extra lengthCan over-shorten if tuned badly
Verifiable rewardsUnit tests, exact match, claim checks against sourcesOnly for tasks that can be checked
KL penalty and early stoppingKeep the policy close to the starting model; stop when human eval stops improvingLimits how far training can move

A gated reward looks like this:

Python
def reward(answer, sources, prompt):    claims = extract_claims(answer)    supported = [nli(sources, c) == "entailed" for c in claims]    if claims and not all(supported):        return 0.0                                     # factual gate: no partial credit    score = 0.6 * follows_instructions(prompt, answer) + 0.4 * helpfulness(answer)    extra = max(0, n_tokens(answer) - target_length(prompt))    return score - 0.001 * extra                      # mild penalty beyond the needed length

The gate removes the trade the model was making: no amount of fluent padding earns reward if a claim is unsupported.

Watch for the signature

Put two lines on one chart: the proxy reward, and accuracy judged by humans (or a trusted gold eval) on a held-out set. When the proxy keeps rising while the gold line flattens or falls, the policy is gaming the reward. Treat that crossing as a stop condition. Also refresh preference data from the current policy's outputs: a reward model trained on last quarter's samples has never seen the new tricks.

A real-life example

Scenario, numbers made up. A legal-help assistant is tuned with a helpfulness reward model. Over training, reward-model score climbs steadily, average answer length grows from 180 to 420 tokens, and a lawyer-graded factual accuracy check falls from 82% to 74%.

The team rebuilds the reward: a claim-level NLI gate against retrieved statutes, length-balanced preference pairs, a small length penalty, a stronger KL penalty, and early stopping on the lawyer-graded set. After retraining, average length is about 210 tokens, factual accuracy is 86%, and users rate answers as helpful as before.

Follow-up questions to expect

  • "Does DPO avoid this?" — No. Direct preference optimisation skips the separate reward model but learns from the same preference pairs, so it can pick up the same length and confidence biases.
  • "What does the KL penalty do?" — It charges the policy for moving far from the reference model, which limits how far it can drift toward strange, high-reward behaviour.
  • "How do you catch new hacks you haven't seen?" — Keep an adversarial slice of confident-but-wrong answers that must score low, and regularly read samples of the highest-reward outputs.