Course Content
Reinforcement Learning from Human Feedback (RLHF)
4 sections · 10 lessons
Preference Data and Reward Models
A team collects 40,000 human ratings for their assistant. They ask annotators to score each response from 1 to 10 on helpfulness. The spreadsheet comes back and the averages look reasonable. Then someone plots the per-annotator distributions.
Annotator A has a mean of 8.4 and a standard deviation of 0.9 — she gives almost everything an 8 or a 9. Annotator B has a mean of 5.1 and uses the full range. Annotator C never uses 1, 2, 9 or 10, so his entire scale is 3 to 8. A response that A scored 8 and B scored 6 is, in each rater's own terms, roughly the same judgement — "solidly above average" — but the numbers disagree by two points. Worse, when the same response was shown to A on a Monday and again three weeks later, she gave it a 7 the second time.
The dataset is not measuring response quality. It is measuring quality confounded with rater generosity, scale usage and mood. Train a regression model on it and you fit the confound as eagerly as the signal.
Now change one thing. Show each annotator two responses side by side and ask a single question: which of these is better? Annotator A's generosity cancels out — she is comparing two responses under the same generous eye. B's harshness cancels out. C's compressed range is irrelevant, because he is not producing a number at all. The comparison isolates the thing you actually want.
Preference data is not a stylistic choice about annotation UI. It is the mechanism that removes per-rater calibration from the signal, and it is why RLHF works at the scale it does.
What preference data actually is
A preference example is a triple: a prompt, a chosen response, and a rejected response.
1{2 "prompt": "My sourdough starter smells like nail polish remover. Is it dead?",3 "chosen": "That acetone smell is usually hunger, not death. It means the yeast has eaten all available flour and is producing acetone-like esters. Pour off any dark liquid on top, discard all but 50g of starter, and feed it 50g flour + 50g water. Repeat every 12 hours. If it's bubbling again within two days, it was fine. If there's pink or orange streaking, that's bacterial contamination and you should start over.",4 "rejected": "It sounds like your starter may have gone bad. Sourdough starters can be tricky! You may want to consider making a new one, or consulting a baking expert for advice on your specific situation."5}The rejected response is not gibberish. That is important. It is fluent, polite, and superficially reasonable — it is exactly the kind of empty, hedge-everything answer a supervised-fine-tuned model produces. The comparison teaches the reward model a distinction that no single demonstration could: specific and actionable beats fluent and vacuous.
Where the candidate pairs come from
You need two responses to compare. There are four common sources, and they have different properties.
| Source of the pair | How it works | Strength | Weakness |
|---|---|---|---|
| Same model, two samples | Sample twice at temperature 0.8–1.0 from the current policy | Both responses are on-distribution; the RM learns exactly the boundary the policy sits on | If the model is uniformly mediocre, both are mediocre and the annotator is picking between two bad options |
| Two model sizes / checkpoints | One from a strong model, one from a weak one | Clear quality gap, high annotator agreement | Too easy — the RM learns to detect the weak model's fingerprint rather than quality |
| Model vs. human-written | Compare a generation against a curated ideal answer | Anchors the top of the scale | Human text has stylistic tells; the RM may learn "sounds human" instead of "is good" |
| Best-of-N | Sample N=4 to 8, have the annotator rank them, expand into pairs | N=4 ranking yields 6 pairs from one annotation session — far more data per minute | Pairs from one session are correlated; treating them as independent overstates your effective sample size |
The best-of-N arithmetic is worth doing explicitly. Ranking 4 responses takes roughly 90 seconds and produces (24)=6 pairs; judging 6 independent pairs takes about 3 minutes. Ranking is twice as fast per pair — but the 6 pairs share responses, so they carry the information of perhaps 3 or 4 independent pairs, not 6.
The tie problem
Roughly 15–25% of comparisons in real datasets are genuine ties — two responses that a careful annotator honestly cannot separate. You have three options, and the choice matters:
- Force a choice. Simple, but you inject coin-flip noise directly into the training labels. Every forced tie is a 50% chance of a wrong gradient.
- Discard ties. Cleanest labels, but you throw away a quarter of your annotation spend, and you bias the dataset toward prompts where quality differences are large — which are the prompts the policy already handles.
- Record ties and use a soft target. Keep them with a label of 0.5 and use a soft-target loss. This is what a proper Bradley-Terry treatment does, and it is the right answer when your annotation tool supports it.
A related trap: a binary label throws away strength of preference — "slightly better" and "dramatically better" become the same example. Collecting a 1–5 strength alongside the choice and weighting the loss by it recovers that information without reintroducing absolute calibration.
Guidelines are the actual product
Annotators do not share your intuitions. "Pick the more helpful response" yields 60–70% agreement between raters — meaning roughly a third of your labels are effectively random relative to each other.
A guideline that works states a priority order and resolves conflicts explicitly:
PRIORITY ORDER (apply in sequence; a higher rule wins outright)1. HARMLESSNESS If exactly one response provides operational instructions for serious harm, the other wins regardless of every other quality.2. TRUTHFULNESS If one response states a verifiable falsehood and the other does not, the truthful one wins - even if it is shorter, blunter, or less well written. A response that says "I'm not certain, but..." is NOT penalised for hedging when the uncertainty is genuine.3. INSTRUCTION COMPLIANCE Did it do what was asked? A brilliant essay when the user asked for three bullet points does not comply.4. HELPFULNESS Actionable and specific beats general and vague. Answers the actual question rather than a nearby easier one.5. STYLE Only reach this rule if 1-4 are genuinely tied.EXPLICIT NON-CRITERIA - do NOT let these influence your choice: * Length. Longer is not better. If two responses say the same thing and one is half as long, the shorter one wins on rule 4. * Formatting. Bullet points are not inherently better than prose. * Agreeableness. A response that tells the user they are wrong, correctly, beats one that agrees to be pleasant.The "explicit non-criteria" block is doing the heaviest lifting on that page. Length bias, formatting bias and agreeableness bias are the three most reliable ways for a reward model to go wrong, and they enter through annotators who were never told not to reward them.
Aggregating multiple raters
Collecting 3–5 judgements per pair costs 3–5 times more, so most datasets label each pair once. Label a subset — 5% is typical — multiple times anyway, because that subset is how you measure whether your guidelines work at all.
The metric is Cohen's kappa, which corrects agreement for chance:
where po is the observed agreement rate and pe is the agreement you would expect from raters choosing at random. On binary A/B choices pe≈0.5, so if two annotators agree on 74% of pairs, κ=(0.74−0.5)/(1−0.5)=0.48. That is moderate agreement at best. Get κ above 0.6 before you scale up; below 0.4 your guidelines are broken and more data will not help.
Low inter-rater agreement is not an annotator problem. It is a specification problem: you have not decided what "better" means precisely enough to be answered consistently.
Scale, cost, and what to spend it on
Real preference dataset sizes, for orientation:
| System | Approximate comparison count | Notes |
|---|---|---|
| Summarisation-from-feedback research | ~64,000 | Narrow task, single domain |
| InstructGPT reward model | ~33,000 prompts, ranked 4–9 ways | Ranking expansion yields far more pairs |
| Llama 2 helpfulness + safety | ~1.4 million | Collected in weekly batches against the current policy |
| Open datasets (Anthropic HH, OASST, UltraFeedback) | ~50,000–350,000 each | Free starting point; quality and domain coverage vary widely |
The practical floor is lower than those numbers suggest: around 5,000 clean, well-covered comparisons train a usable domain-specific reward model, and 50,000 gets a solid general-purpose one. Below roughly 2,000 the model overfits before it learns anything general.
Coverage matters more than raw count. If 70% of your prompts are coding questions, your reward model is a coding-quality scorer that emits confident nonsense on medical questions. Build a prompt taxonomy first and sample against target proportions.
Getting more signal per unit of spend
- Active learning. Score candidate pairs with your current reward model and send annotators the pairs where the predicted preference is closest to 0.5. Those are the pairs the model is most uncertain about, and therefore the ones with the most information. This usually reaches a given accuracy with fewer labels than sampling pairs uniformly. The catch: uncertainty sampling systematically over-selects genuinely ambiguous pairs, so mix in 20–30% uniformly sampled pairs or your dataset becomes all edge cases.
- Model-assisted pre-labelling. Have a strong model pre-select a winner and ask humans to confirm or override; confirming is faster than judging from scratch. The catch is anchoring — humans agree with the pre-label more often than they would have chosen it. Blind a control fraction to measure the effect.
- Synthetic preferences. Use a capable model as the judge, generating preference labels directly. Cheap and fast. Published studies of strong judge models report agreement with human labels of roughly 65–85% depending on the task and on whether ties count (Zheng et al., 2023, for GPT-4 as a judge; Lee et al., 2023, for AI-labelled summarisation preferences). Agreement is high on clear cases and much worse on subtle ones. Reasonable for bulk coverage, dangerous as your only source — you inherit the judge model's biases wholesale, including its position bias (a judge shown A first prefers A more often, an effect large enough that you must randomise and average over both orderings).
- Iterative collection. Collect a batch, train, deploy, then collect the next batch on the new model's outputs. This keeps the reward model's training distribution matched to the policy it scores.
Building the reward model
A reward model is a language model with its next-token prediction head removed and replaced by a single scalar output. You feed it the prompt and response concatenated; it returns one number.
1import torch2import torch.nn as nn3from transformers import AutoModel, AutoTokenizer45class RewardModel(nn.Module):6 def __init__(self, base_name="microsoft/deberta-v3-base"):7 super().__init__()8 self.backbone = AutoModel.from_pretrained(base_name)9 hidden = self.backbone.config.hidden_size10 self.head = nn.Linear(hidden, 1)11 # Small init keeps early rewards near zero, which stops the12 # value function from chasing a wild scale on step one.13 nn.init.normal_(self.head.weight, std=1e-3)14 nn.init.zeros_(self.head.bias)1516 def forward(self, input_ids, attention_mask):17 out = self.backbone(input_ids=input_ids,18 attention_mask=attention_mask)19 hidden = out.last_hidden_state # (B, T, H)20 # Pool at the LAST NON-PAD token, not position -1, and not21 # a mean over the sequence. See the note below.22 last_idx = attention_mask.sum(dim=1) - 1 # (B,)23 pooled = hidden[torch.arange(hidden.size(0)), last_idx]24 return self.head(pooled).squeeze(-1) # (B,)That pooling detail causes more silent failures than any other line in a reward model. Index position -1 on a right-padded batch and you read a padding token's hidden state, making reward a function of how much padding the example happened to get — the loss still falls, but the model has learned batch composition, not quality. Mean-pool instead and you dilute the response signal with the prompt's.
Three architectural options
| Design | Output | Best for | Cost |
|---|---|---|---|
| Sequence classification — one scalar for the whole response | r(x,y)∈R | The default. Matches how humans judged (whole response vs whole response) | Cheapest; one forward pass |
| Token-level rewards — a scalar per token | r(x,y1:t) for each t | Fine-grained credit assignment: which sentence caused the bad rating? | Needs span-level annotation, which is far more expensive |
| Multi-head — separate scalars per attribute | helpfulness, harmlessness, honesty | When you need to trade off objectives explicitly, or apply a hard safety veto | Needs preferences labelled per attribute; one shared backbone keeps compute modest |
The multi-head design solves a real problem. With a single scalar, a helpful-but-slightly-unsafe response can outscore a safe-but-bland one, and the scalar has already collapsed the trade-off beyond recovery. With separate heads you combine explicitly — r=0.7rhelp+0.3rsafe — and can then veto: if rsafe falls below −2, return a large negative reward regardless of helpfulness. A single-scalar model cannot express that veto.
Choosing the backbone
The reward model need not match the policy in size, but it must be large enough to understand what it scores. A 400M-parameter encoder can tell a rambling answer from a crisp one; it cannot tell a correct proof from a subtly wrong one. Common practice is between one-tenth of the policy's size and full parity, initialised from the SFT checkpoint so it starts with the same task understanding.
The training objective, worked through
The loss follows directly from the Bradley-Terry model of paired comparisons, which says the probability a human prefers yw over yl is a sigmoid of the score gap:
Read it as: for each labelled pair, push the winner's score above the loser's, and stop pushing once the gap is comfortable. Only the difference appears, so the absolute scale is arbitrary — adding 5 to every score leaves the loss unchanged. A reward of 3.2 is meaningless alone; it means something only against other scores from the same model.
Work the numbers on one batch of three pairs:
| Pair | r(yw) | r(yl) | Gap | σ(gap) | Loss −logσ |
|---|---|---|---|---|---|
| 1 — ranked correctly, confident | 2.10 | 0.30 | 1.80 | 0.858 | 0.153 |
| 2 — ranked correctly, barely | 0.90 | 0.80 | 0.10 | 0.525 | 0.644 |
| 3 — ranked backwards | -0.40 | 1.10 | -1.50 | 0.182 | 1.701 |
Batch loss is (0.153+0.644+1.701)/3=0.833. Notice how the gradient budget is allocated: pair 3 contributes over eleven times the loss of pair 1. The objective automatically concentrates on the examples it currently gets wrong, and stops wasting capacity on pairs it already has right. That is the practical consequence of the sigmoid's saturating shape.
The implementation is four lines:
1import torch.nn.functional as F23def reward_loss(model, chosen_ids, chosen_mask, rejected_ids, rejected_mask):4 r_chosen = model(chosen_ids, chosen_mask) # (B,)5 r_rejected = model(rejected_ids, rejected_mask) # (B,)6 # -log(sigmoid(gap)) == softplus(-gap), numerically stable7 loss = F.softplus(-(r_chosen - r_rejected)).mean()8 accuracy = (r_chosen > r_rejected).float().mean()9 return loss, accuracy, r_chosen.mean(), r_rejected.mean()Two implementation notes matter. Use softplus(-gap) rather than -log(sigmoid(gap)), which underflows to -inf when the gap is strongly negative early in training. And keep the chosen and rejected sequences of a pair in the same forward batch — split across gradient-accumulation steps, the difference is computed on stale activations and the gradient is wrong.
Evaluating a reward model
| Metric | What it measures | Healthy range | What it misses |
|---|---|---|---|
| Pairwise accuracy | Fraction of held-out pairs where r(yw)>r(yl) | 0.65–0.75 | Says nothing about calibration or about behaviour off-distribution |
| Reward gap distribution | Histogram of r(yw)−r(yl) | Centred positive, moderate spread | — |
| Spearman correlation | Rank correlation between model scores and human ratings on a scored set | 0.5–0.7 | Requires a separately collected absolute-rating set |
| AUC-ROC | Ranking quality independent of any threshold | 0.7–0.85 | Insensitive to how badly the wrong pairs are wrong |
| Length correlation | Correlation between r and token count | Below about 0.2 | The single most important diagnostic, and the one most often skipped |
The accuracy ceiling surprises people. A reward model that scores 0.68 on held-out pairs sounds mediocre until you check the human ceiling: on the same data, two independent annotators agree only about 72–75% of the time (InstructGPT reported about 73% agreement among its labelers). The reward model is close to the noise floor of the labels themselves. Chasing 0.90 accuracy is not a sign of a better model; it is a sign that your validation split leaked, or that your pairs are trivially easy, or that the model has found a shortcut.
Public benchmarks give you a second, outside view. RewardBench (Lambert et al., 2024) and its 2025 successor score reward models on curated chosen/rejected pairs across chat, safety and reasoning. They are useful for choosing an off-the-shelf reward model to start from. They do not replace an evaluation on your own prompts, because a reward model can rank well on a public benchmark and still be poor at your domain.
If your reward model's held-out accuracy exceeds your inter-annotator agreement rate, it has learned something your annotators did not agree on. That is a bug, not an achievement.
Splitting the data correctly
Split by prompt, never by pair. If "explain gradient descent" appears in training with one pair and in validation with another, the model has seen the prompt and can memorise the shape of a good answer. Group all pairs sharing a prompt into one split — typically 90% train, 5% validation, 5% test touched once. Also hold out whole prompt categories to measure out-of-domain generalisation, the failure mode you actually hit in production.
The three ways reward models go wrong
Reward hacking
The reward model has systematic biases, and the policy trained against it will find them. Concrete, documented cases:
- Length. Annotators mildly favour thorough answers. The RM learns a positive length coefficient. The policy pads. Mean response length climbs from 180 to 600 tokens over training; reward rises steadily; blind human evaluation of the final model rates it worse than the starting point.
- Format mimicry. Bulleted answers scored slightly higher in the data, so the policy returns a three-section bulleted report to "what is the capital of Peru?"
- Confidence theatre. Hedged answers scored lower, so the policy drops hedges — including on questions where it is genuinely uncertain. Reward up, honesty down.
- Sycophancy. Agreement scored higher, so the policy reverses a correct answer when the user pushes back with "are you sure?"
- Keyword stuffing. If safety-flavoured phrasing correlated with high scores, the policy learns to sprinkle "it's important to consider" into unrelated answers.
Detection: track mean length, refusal rate, formatting-marker frequency and repeated-phrase rate on every evaluation, alongside reward. A reward curve that rises while any surface statistic moves sharply is the signature of a hack.
Overfitting
Reward models overfit fast — often within a single epoch on datasets under 50,000 pairs. The tell is validation accuracy plateauing or dipping while the reward gap on training data keeps widening: the model is becoming more confident about things it already had right rather than learning anything new. Defences: train for one epoch, use a low learning rate (1e-5 to 5e-6), apply dropout in the backbone, and consider ensembling three reward models with different seeds and taking the minimum score (Coste et al., 2023, found such conservative ensembles reduce over-optimisation) — the minimum is conservative exactly where the ensemble disagrees, which is exactly where hacking happens.
Distribution shift
This is the subtle one. Your reward model was trained on responses from the SFT model. During policy optimisation, the policy moves. After a few thousand steps it is producing text unlike anything in the reward model's training set — and the reward model has no idea it is out of its depth. It confidently assigns high scores to text it has never seen the like of.
Two defences. First, the KL penalty against a frozen reference model, which keeps the policy near the distribution the reward model understands. Second, iterative data collection: after each round of policy training, sample from the new policy, collect fresh comparisons on those samples, and retrain the reward model. This is why production pipelines run in weekly cycles rather than as a single pass.
What to do first when you build one
The order of operations is not obvious, and getting it wrong wastes months.
Write the guidelines before collecting anything. Then have three people label 200 pairs against them and compute kappa. If it is under 0.6, the guidelines are ambiguous — find the pairs they disagreed on, work out which rule failed to resolve them, and rewrite. Doing this costs two days and saves you from a 40,000-example dataset that encodes disagreement.
Build the length diagnostic before the model. On day one, plot reward against token count on your validation set. If the correlation exceeds 0.3, stop and fix the data — either your guidelines did not forbid length bias or your candidate generation systematically pairs long with short. Everything downstream inherits this.
Hold out a category, not just a random sample. Random held-out pairs tell you the model generalises to new pairs from the same distribution. They tell you nothing about the medical question your users will ask on launch day. Reserve an entire prompt category and check accuracy on it separately; a drop from 0.71 to 0.54 on unseen domains is normal and worth knowing about before deployment.
Expect to collect data more than once. The largest structural difference between pipelines that produce a nice curve and pipelines that produce a good model is that the latter collect preferences iteratively against the current policy. Budget for three or four rounds, not one.
Treat the reward model as an instrument with a calibrated range. It is accurate near the distribution it was trained on and unreliable outside it. Every practice above — the KL leash, iterative collection, ensembling, off-domain evaluation — exists either to widen that range or to stop the policy leaving it.