Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your reward model prefers 'helpful' answers. The LLM learns to sound confident and verbose — but becomes factually worse. How do you design reward signals that don't get gamed?
What you need to know
How a reward gets gamed
In RLHF (reinforcement learning from human feedback), people compare pairs of answers, a reward model learns to predict which one they prefer, and the LLM is trained to score high on that reward model. The reward model is a proxy for "good". If long, confident answers were usually preferred, the reward model learns "long and confident is good", and the LLM happily produces that — even when it is wrong.
Length bias is one of the most widely reported cases: preference models trained without controls tend to reward longer answers.
Four defences
| Defence | What it does | Trade-off |
|---|---|---|
| Decomposed reward with a gate | Separate scores; wrong answers get zero | More labelling and more models to maintain |
| Length control | Balance pairs so the chosen answer is not usually longer, or penalise extra length | Can over-shorten if tuned badly |
| Verifiable rewards | Unit tests, exact match, claim checks against sources | Only for tasks that can be checked |
| KL penalty and early stopping | Keep the policy close to the starting model; stop when human eval stops improving | Limits how far training can move |
A gated reward looks like this:
1def reward(answer, sources, prompt):2 claims = extract_claims(answer)3 supported = [nli(sources, c) == "entailed" for c in claims]4 if claims and not all(supported):5 return 0.0 # factual gate: no partial credit6 score = 0.6 * follows_instructions(prompt, answer) + 0.4 * helpfulness(answer)7 extra = max(0, n_tokens(answer) - target_length(prompt))8 return score - 0.001 * extra # mild penalty beyond the needed lengthThe gate removes the trade the model was making: no amount of fluent padding earns reward if a claim is unsupported.
Watch for the signature
Put two lines on one chart: the proxy reward, and accuracy judged by humans (or a trusted gold eval) on a held-out set. When the proxy keeps rising while the gold line flattens or falls, the policy is gaming the reward. Treat that crossing as a stop condition. Also refresh preference data from the current policy's outputs: a reward model trained on last quarter's samples has never seen the new tricks.
A real-life example
Scenario, numbers made up. A legal-help assistant is tuned with a helpfulness reward model. Over training, reward-model score climbs steadily, average answer length grows from 180 to 420 tokens, and a lawyer-graded factual accuracy check falls from 82% to 74%.
The team rebuilds the reward: a claim-level NLI gate against retrieved statutes, length-balanced preference pairs, a small length penalty, a stronger KL penalty, and early stopping on the lawyer-graded set. After retraining, average length is about 210 tokens, factual accuracy is 86%, and users rate answers as helpful as before.
Follow-up questions to expect
- "Does DPO avoid this?" — No. Direct preference optimisation skips the separate reward model but learns from the same preference pairs, so it can pick up the same length and confidence biases.
- "What does the KL penalty do?" — It charges the policy for moving far from the reference model, which limits how far it can drift toward strange, high-reward behaviour.
- "How do you catch new hacks you haven't seen?" — Keep an adversarial slice of confident-but-wrong answers that must score low, and regularly read samples of the highest-reward outputs.