Course Content
Reinforcement Learning from Human Feedback (RLHF)
4 sections · 10 lessons
Training Reward Models and PPO Fine-tuning
A reward model finishes training with 98.4% accuracy on the held-out set. That should be impossible — two human annotators only agree with each other about 74% of the time, so a model that beats them by 24 points has not learned human preference. It has learned something else.
The bug takes an hour to find. The tokeniser was called separately on the chosen and rejected texts, with max_length=512 and truncation=True. Rejected responses in this dataset happened to be longer on average, so 31% of them were cut off mid-word. The reward model learned a rule of stunning simplicity: if the text ends mid-sentence, score it low. On held-out data drawn from the same pipeline, that rule was right 98% of the time.
Then the policy trained against it. Reward climbed beautifully. The policy's discovery was equally simple: never stop mid-sentence, and pad the ending with an unambiguous closing flourish. Every response now ends with three variations of "I hope this helps! Let me know if you have any other questions!"
A reward model's accuracy tells you how well it separates your chosen and rejected sets. It does not tell you that the thing separating them is quality.
What follows is the pipeline done properly, with the checks that would have caught that in the first ten minutes.
The shape of the whole thing
preference pairs prompts only (prompt, chosen, rejected) (no responses needed) | | v | [1] TRAIN REWARD MODEL | base + scalar head | Bradley-Terry pairwise loss | output: r(x, y) -> one float | | | +-------------> frozen -------------+ | v [2] PPO POLICY TRAINING policy (trainable, init = SFT) value (trainable) ref (frozen copy of SFT) reward (frozen, from step 1) | v aligned modelTwo stages, four models, and a strict rule: the reward model must be frozen during stage 2. If you train it concurrently on the policy's own outputs, the two models chase each other and the reward scale diverges within a few hundred steps.
Environment
1pip install "trl>=1.0,<2" "transformers>=5" peft datasets \2 accelerate bitsandbytes wandb scipy34# The code in this lesson was checked against trl 1.13 and5# transformers 5.17. Pin the versions you develop against and read6# the docs for that exact version.7pip freeze | grep -E "^(trl|transformers|peft)=="That comment is not boilerplate. The reinforcement-learning trainers in this ecosystem have had breaking changes between releases more than once; TRL's PPO trainer was removed outright. Code copied from a blog post written a year earlier frequently fails with an import or argument error, so pin your versions and treat the installed docs as authoritative.
Stage 1: the reward model
Data preparation, with the traps closed
1from datasets import load_dataset2from transformers import AutoTokenizer34MODEL = "distilroberta-base" # small enough to iterate quickly5MAXLEN = 51267tok = AutoTokenizer.from_pretrained(MODEL)8raw = load_dataset("Anthropic/hh-rlhf", split="train[:20000]")910# hh-rlhf stores the whole conversation in "chosen" and "rejected";11# RewardTrainer accepts these text columns as they are.1213# TRAP 1 + 2: never let one side be cut and the other not. Drop pairs14# where EITHER side is too long, so truncation cannot become a15# learnable signal. If you lose more than ~15%, raise MAXLEN instead.16def fits(ex):17 return (len(tok(ex["chosen"])["input_ids"]) <= MAXLEN and18 len(tok(ex["rejected"])["input_ids"]) <= MAXLEN)1920ds = raw.filter(fits)21print(f"dropped {100*(len(raw)-len(ds))/len(raw):.1f}% as too long")2223# TRAP 3: split by PROMPT, never by pair. Pairs sharing a prompt in24# both train and eval let the model memorise good answers per prompt.25ds = ds.train_test_split(test_size=0.05, seed=0)Three traps, all of which silently inflate accuracy. Truncation asymmetry is the one from the opening. Prompt leakage is the subtler cousin: with 4 responses ranked per prompt you get 6 pairs, and a random split scatters them across train and test, so the model has already seen the prompt and most of its answers.
Current TRL RewardTrainer also drops, rather than truncates, pairs longer than max_length, but measure the loss yourself: if it removes a third of your data, you want to know. And train_test_split splits by row. With one pair per prompt, as in hh-rlhf, that is a split by prompt; with several pairs per prompt you must group by prompt first.
The model
1from transformers import AutoModelForSequenceClassification23rm = AutoModelForSequenceClassification.from_pretrained(4 MODEL,5 num_labels=1, # ONE output: a scalar score6 problem_type="regression",7)8# The head is randomly initialised. Keep it small so early rewards9# start near zero rather than at some arbitrary large scale.10rm.classifier.out_proj.weight.data.normal_(std=1e-3)11rm.classifier.out_proj.bias.data.zero_()num_labels=1 is what turns a classifier into a reward model. There is no softmax and no class; the forward pass returns a single unbounded real number per sequence. Anything above zero is not "good" and anything below is not "bad" — only differences between scores carry meaning, because the training objective sees only differences.
Why this loss, specifically
The training data is comparisons, not scores, so you need an objective that turns "A beat B" into pressure on two numbers. The Bradley-Terry model, borrowed from the statistics of paired contests, supplies it. It assumes the chance of preferring A is a sigmoid of the gap between their latent strengths:
Fitting by maximum likelihood gives the loss:
In plain terms: for every labelled pair, raise the winner's score above the loser's — and ease off once the gap is comfortable.
The shape earns its place for three reasons. It depends only on the difference, which is exactly the information a comparison carries and nothing more, so the model is never asked to invent an absolute scale that the data cannot support. It saturates, so once a pair is ranked correctly with a healthy gap the gradient nearly vanishes and capacity moves to pairs still getting it wrong. And it is smooth and convex in the gap, so optimisation is well behaved — unlike a hinge or a hard ranking constraint.
Work a batch of three pairs to see the gradient allocation:
| Pair | r(yw) | r(yl) | Gap d | σ(d) | Loss −logσ(d) | ∂L/∂d=σ(d)−1 |
|---|---|---|---|---|---|---|
| Confidently right | 2.10 | 0.30 | 1.80 | 0.858 | 0.153 | -0.142 |
| Barely right | 0.90 | 0.80 | 0.10 | 0.525 | 0.644 | -0.475 |
| Wrong | -0.40 | 1.10 | -1.50 | 0.182 | 1.701 | -0.818 |
Mean loss is (0.153+0.644+1.701)/3=0.833. The gradient magnitudes tell the real story: the wrongly ranked pair pushes 5.8 times harder than the confidently correct one. The objective self-prioritises without you writing any curriculum logic.
Training it
1from trl import RewardTrainer, RewardConfig23cfg = RewardConfig(4 output_dir="rm-out",5 per_device_train_batch_size=8,6 gradient_accumulation_steps=4, # effective batch 327 num_train_epochs=1, # RMs overfit fast - start here8 learning_rate=1e-5,9 warmup_ratio=0.03,10 lr_scheduler_type="cosine",11 max_length=MAXLEN,12 eval_strategy="steps",13 eval_steps=100,14 bf16=True,15)1617trainer = RewardTrainer(18 model=rm, args=cfg, processing_class=tok,19 train_dataset=ds["train"], # text columns: chosen, rejected20 eval_dataset=ds["test"],21)22trainer.train()One epoch is deliberate, and it has to be set explicitly: RewardConfig defaults to three epochs and a learning rate of 1e-4. On datasets below roughly 50,000 pairs, reward models begin memorising in the second epoch; the tell is training reward gaps widening while eval accuracy plateaus or slips. If you want the loss written out rather than delegated:
1import torch.nn.functional as F23def pairwise_loss(model, batch):4 r_w = model(input_ids=batch["input_ids_chosen"],5 attention_mask=batch["attention_mask_chosen"]).logits.squeeze(-1)6 r_l = model(input_ids=batch["input_ids_rejected"],7 attention_mask=batch["attention_mask_rejected"]).logits.squeeze(-1)8 # softplus(-d) == -log(sigmoid(d)), but does not underflow to -inf9 loss = F.softplus(-(r_w - r_l)).mean()10 return loss, (r_w > r_l).float().mean(), (r_w - r_l).mean()Evaluating it — three checks, not one
1import numpy as np, torch2from scipy.stats import pearsonr34@torch.no_grad()5def audit(model, tok, eval_ds, device="cuda"):6 model.eval()7 gaps, lens_w, lens_l, scores = [], [], [], []8 for ex in eval_ds:9 rw = score(model, tok, ex["chosen"], device)10 rl = score(model, tok, ex["rejected"], device)11 gaps.append(rw - rl)12 lens_w.append(len(tok(ex["chosen"])["input_ids"]))13 lens_l.append(len(tok(ex["rejected"])["input_ids"]))14 scores += [rw, rl]15 gaps = np.array(gaps)16 lens = np.array(lens_w + lens_l)1718 acc = float((gaps > 0).mean())19 len_corr = float(pearsonr(scores, lens)[0])20 # If you always picked the LONGER response, how often would you win?21 len_base = float((np.array(lens_w) > np.array(lens_l)).mean())2223 print(f"pairwise accuracy {acc:.3f}")24 print(f"length baseline {len_base:.3f} <- beat this by a lot")25 print(f"reward~length corr {len_corr:+.3f} <- keep below 0.20")26 print(f"mean gap {gaps.mean():+.3f}")27 return acc, len_base, len_corrThe length baseline is the check that would have saved the team in the opening story, and almost nobody computes it. Suppose your reward model hits 0.71 accuracy. Impressive — unless a policy of "always pick the longer response" also scores 0.68 on the same data, in which case your model has contributed three points over a heuristic that requires no learning at all. Healthy numbers look like: accuracy 0.70, length baseline 0.55, length correlation 0.12.
| Symptom | Likely cause | Fix |
|---|---|---|
| Accuracy above 0.90 | Leakage, truncation artefact, or trivially easy pairs | Check the length baseline; verify split is by prompt; inspect 20 pairs by hand |
| Accuracy near 0.50 | Label noise, wrong pooling, or chosen/rejected swapped | Verify labels on 30 examples; check the pooling index; print score pairs |
| Length correlation above 0.35 | Guidelines did not forbid length bias, or asymmetric generation | Fix at the data layer; length-match a subset and retrain |
| Eval accuracy falls after step N | Overfitting | Stop at N; reduce to one epoch; lower the learning rate |
| Reward scale drifts to ±40 | Head initialised too large, or learning rate too high | Small head init; the scale is arbitrary but a wild one destabilises the value model later |
Scoring text with it
1@torch.no_grad()2def score(model, tok, text, device="cuda"):3 enc = tok(text, return_tensors="pt", truncation=True,4 max_length=MAXLEN).to(device)5 return model(**enc).logits.squeeze().item()67prompt = "Human: My laptop won't turn on. Assistant:"8candidates = [9 " Have you tried turning it off and on again?",10 (" Let's isolate it. 1) Hold power for 15s to force a drain, then"11 " plug in and retry. 2) Check the charger LED - no light means"12 " charger or port. 3) Try a known-good charger. If the fan spins"13 " but the screen stays dark, it's likely display or GPU."),14 " I'm sorry to hear that. Computers can be frustrating sometimes!",15]16for c in candidates:17 print(f"{score(rm, tok, prompt + c):+.3f} {c[:52]}...")1819# Typical output:20# -0.412 Have you tried turning it off and on again?...21# +1.847 Let's isolate it. 1) Hold power for 15s to force...22# -0.938 I'm sorry to hear that. Computers can be frustra...Do this on 20 hand-written candidates before you start PPO. It takes fifteen minutes and it is the only way to see whether the model has learned quality or a proxy. If the empathetic-but-useless response outscores the actionable one, your policy will learn to be empathetic and useless — with great efficiency.
Score twenty responses you wrote yourself and read the ordering. If the reward model disagrees with you there, every hour you then spend on PPO is spent making the policy better at being wrong.
Stage 2: what PPO is actually doing here
The reward model gives a score for a whole response. Gradient descent needs a per-token signal. PPO bridges that gap with three ideas.
Advantages instead of raw rewards. A value model predicts the expected return from each partial response, and the advantage is the difference between what happened and what was expected. This centres the signal: on an easy prompt where every response scores near 8.0, raw rewards say "all good" while advantages still separate the 8.4 from the 7.8.
A clipped objective so one batch can be reused. Generating rollouts is far more expensive than the gradient step, so you want several optimisation passes per rollout. But policy gradients are only valid on-policy. PPO corrects with an importance ratio and then bounds it:
Plainly: score the update both with the true importance ratio and with the ratio pinned inside [0.8,1.2] , then keep the smaller. Taking the minimum removes any reward for pushing a token further once it has already moved a lot — while still allowing the full gradient to pull a token back if the policy moved it the wrong way.
A KL penalty folded into the reward. Clipping bounds each step but not cumulative drift. So every token pays a tax for how far the policy has moved from a frozen reference:
Reading it: charge a small penalty at every token proportional to the local drift, and add the reward model's score only at the final token. Spreading the penalty across tokens rather than applying it in one lump at the end is what lets the value model attribute drift to the tokens that caused it.
Running PPO
Older tutorials run this stage with TRL's PPOTrainer and a ppo_trainer.step() loop. That code no longer runs: TRL 1.x no longer ships PPOTrainer or AutoModelForCausalLMWithValueHead, and its online trainers today are the critic-free GRPOTrainer and RLOOTrainer (section 3 covers them). For PPO with a critic and a learned reward model, current practice is a framework built around rollout generation: OpenRLHF or veRL. Both generate with vLLM, train with DeepSpeed or FSDP, and use Ray to place the four models on GPUs.
An OpenRLHF launch, trimmed to the flags that map onto this lesson. The flag names are from the project's README in 2026 and have been renamed between releases, so check them against the version you install.
1pip install "openrlhf[vllm]"23ray start --head --num-gpus 84ray job submit --address="http://127.0.0.1:8265" \5 -- python3 -m openrlhf.cli.train_ppo_ray \6 --actor.model_name_or_path ./sft-model \7 --reward.model_name_or_path ./rm-out \8 --data.prompt_dataset ./prompts \9 --data.input_key prompt \10 --data.apply_chat_template \11 --rollout.batch_size 1024 \12 --rollout.n_samples_per_prompt 1 \13 --train.batch_size 128 \14 --train.max_epochs 1 \15 --generate_max_len 1024 \16 --actor.adam.lr 5e-7 \17 --critic.adam.lr 9e-6 \18 --algo.kl.init_coef 0.01 \19 --reward.normalize_enable \20 --ds.zero_stage 3 \21 --ds.param_dtype bf16Map the flags back to the algorithm. --algo.kl.init_coef is β. --reward.normalize_enable is the reward normalisation described below. The critic has its own, higher learning rate. Leaving --algo.advantage.estimator unset gives PPO with GAE; setting it to rloo or group_norm switches to the critic-free methods in section 3 without changing anything else. OpenRLHF expects a reward model in its own format, so train it with its openrlhf.cli.train_rm script, or pass --reward.remote_url pointing at a Python reward function instead.
veRL expresses the same run as python3 -m verl.trainer.main_ppo with Hydra-style overrides such as algorithm.adv_estimator=gae and algorithm.kl_ctrl.kl_coef=0.001. Its quickstart trains Qwen2.5-0.5B-Instruct on GSM8K on a single GPU, which makes it a good first PPO run.
Whichever framework you use, one iteration does the same thing. Written out:
1logp_old = logprobs_of(policy, queries, responses) # AT GENERATION TIME2logp_ref = logprobs_of(ref, queries, responses) # frozen model3values = value_head(policy, queries, responses)45# per-token reward: KL tax everywhere, RM score at the last token6kl = logp_old - logp_ref7per_tok_r = -beta * kl8per_tok_r[:, -1] += rm_scores910# GAE11adv, returns = [], []12last = 0.013for t in reversed(range(T)):14 next_v = values[:, t + 1] if t < T - 1 else 0.015 delta = per_tok_r[:, t] + gamma * next_v - values[:, t]16 last = delta + gamma * lam * last17 adv.insert(0, last)18adv = torch.stack(adv, dim=1)19returns = adv + values20adv = (adv - adv.mean()) / (adv.std() + 1e-8) # per-batch whitening2122for _ in range(ppo_epochs):23 for mb in minibatches(...):24 logp_new = logprobs_of(policy, mb.queries, mb.responses)25 rho = torch.exp(logp_new - mb.logp_old) # NOT recomputed26 pg = -torch.min(rho * mb.adv,27 rho.clamp(1 - eps, 1 + eps) * mb.adv).mean()28 v_clipped = mb.values + (value_head(...) - mb.values).clamp(-cv, cv)29 vf = 0.5 * torch.max((value_head(...) - mb.returns) ** 2,30 (v_clipped - mb.returns) ** 2).mean()31 loss = pg + vf_coef * vf32 loss.backward()33 torch.nn.utils.clip_grad_norm_(policy.parameters(), 1.0)34 opt.step(); opt.zero_grad()The comment on logp_old marks the most damaging bug in PPO implementations. Those log-probabilities must be the ones recorded when the tokens were generated. Recompute them from the current policy and every ratio is exactly 1.0 at the first inner epoch, clipping never engages, and you are running unconstrained policy gradient while your logs cheerfully report a clip fraction of zero.
Stability techniques and the failure each one prevents
| Technique | Failure it prevents | How |
|---|---|---|
| Reward normalisation (running mean/std) | Value model divergence | The RM's raw scale is arbitrary and may drift as the policy moves. A critic regressing on an unbounded moving target diverges; normalising to roughly zero mean, unit variance keeps it in range |
| Advantage whitening per batch | Wildly varying step sizes | Puts advantages on a consistent scale so a batch of easy prompts does not produce a giant update |
| Adaptive KL coefficient | KL explosion or total stagnation | Measures actual KL each step; if above target, raise β; if below, lower it. Turns a brittle constant into a controller |
| Value clipping | Critic overshoot | Bounds how far the value prediction may move from the rollout-time value, mirroring the policy clip |
| Gradient norm clipping (1.0) | Single-batch catastrophe | One pathological batch cannot produce an update large enough to wreck the model |
| Sampling with varied lengths and temperature 1.0 | Learning stalls; length collapse | Without generation diversity all responses in a batch are near-identical, advantages are near-zero, and there is nothing to learn from |
| Reward-model ensemble, take the minimum | Reward hacking | Exploits usually fool one model, not three. The minimum is pessimistic precisely where the ensemble disagrees |
Worked example of why reward normalisation matters. Suppose the RM outputs scores in the range [−3,+12] with mean 4.2 and standard deviation 3.1. The value head is randomly initialised and starts predicting near 0. Its first regression targets are around 4.2, giving early advantages of roughly +4 on essentially every token regardless of quality — so the first few hundred steps uniformly increase the probability of whatever the policy happened to sample. Normalise, and those same early advantages centre near zero, carrying only the relative information you actually want.
Every stability technique in a PPO implementation exists because someone shipped a run without it and watched a model destroy itself. None of them are optional decorations.
Monitoring: the six numbers to watch
| Metric | Healthy | Alarm | Reading |
|---|---|---|---|
| Mean reward | Rises steadily then plateaus | Rises without limit | Unbounded growth is hacking, not learning |
| KL to reference | Settles near your target, 4–10 | Past 30, or stuck below 1 | Too high: degenerate text. Too low: you trained nothing |
| Policy clip fraction | 0.05 – 0.25 | Above 0.4 | Too much clipping means steps too large or too many inner epochs |
| Value loss | Falls then flattens | Rising | Value model diverging — normalise rewards, lower its LR |
| Mean response length | Roughly stable | Climbing steadily | The single clearest fingerprint of length hacking |
| Policy entropy | Declines gently | Falls off a cliff | Mode collapse; add an entropy bonus and raise β |
Each framework logs these under its own names (veRL, for example, uses actor/pg_clipfrac, critic/vf_loss and response_length/mean). Read them together, never alone. Reward up with length stable and KL near target is real progress. Reward up with length climbing and entropy falling is a model learning to exploit your scorer, and the reward curve will look identical in both cases.
Pitfalls worth naming
- Reward model and policy on different tokenisers. Perfectly legal — they are separate models — but you must decode the policy's tokens to text and re-encode with the reward model's tokeniser. Feeding policy token IDs straight into the reward model produces scores from nonsense text, and nothing errors.
- Special tokens leaking into the score. If the policy emits padding or an unexpected chat template token, the reward model has never seen that pattern and its score is arbitrary. Strip and normalise before scoring.
- Greedy decoding during rollouts. Zero diversity, zero advantage spread, no learning. The training loop runs happily and the reward stays flat.
- Forgetting to freeze the reference. If the reference model shares parameters with the policy, KL is identically zero, the leash is off, and you find out three hours in.
- Training more than one epoch on a small reward-model dataset. The reward model memorises, and the policy then optimises against memorised quirks.
- Judging success from the reward curve. Hold out prompts that never touch training and run a blind side-by-side against the SFT model every few hundred steps. That number, not the reward, decides whether you ship.
How to actually run this the first time
Debug the entire pipeline at toy scale before touching a GPU cluster. A 125M policy, a 125M reward model, 500 preference pairs, 50 PPO steps. Everything that will break structurally — tokeniser mismatch, reward shape errors, frozen-model mix-ups, KL identically zero — breaks here, in minutes, on one GPU. Scaling a working pipeline is routine; debugging a broken one at 7B is not.
Audit the reward model by hand before it trains anything. Twenty candidate responses you wrote yourself, scored and sorted. If the ordering matches your judgement, proceed. If it does not, no amount of PPO tuning will help, because the policy will faithfully optimise toward whatever the reward model actually believes.
Set an adaptive KL target and let it do the work. A target of 6 total nats per response is a sane default. Fixing β by hand means retuning it every time the reward scale or response length changes; the controller absorbs that automatically.
Checkpoint every 100 steps and keep a manual sample log. Save five generations for the same fixed prompts at every checkpoint, in a plain text file you actually read. Degradation is obvious to a human two hundred steps before it is obvious in any metric, and the checkpoint you want is always the one just before you noticed.
Stop earlier than feels right. The best checkpoint by human evaluation is very often not the one with the highest reward. Training until the reward plateaus reliably overshoots into the region where the policy has begun exploiting the reward model rather than satisfying it.