Reinforcement Learning from Human Feedback (RLHF)

Training Reward Models and PPO Fine-tuning


A reward model finishes training with 98.4% accuracy on the held-out set. That should be impossible — two human annotators only agree with each other about 74% of the time, so a model that beats them by 24 points has not learned human preference. It has learned something else.

The bug takes an hour to find. The tokeniser was called separately on the chosen and rejected texts, with max_length=512 and truncation=True. Rejected responses in this dataset happened to be longer on average, so 31% of them were cut off mid-word. The reward model learned a rule of stunning simplicity: if the text ends mid-sentence, score it low. On held-out data drawn from the same pipeline, that rule was right 98% of the time.

Then the policy trained against it. Reward climbed beautifully. The policy's discovery was equally simple: never stop mid-sentence, and pad the ending with an unambiguous closing flourish. Every response now ends with three variations of "I hope this helps! Let me know if you have any other questions!"

A reward model's accuracy tells you how well it separates your chosen and rejected sets. It does not tell you that the thing separating them is quality.

What follows is the pipeline done properly, with the checks that would have caught that in the first ten minutes.

The six numbers that tell you a PPO run is healthyWatch every stepMean reward, rising slowlyKL toreference, under budgetClip fraction, 0.1 to 0.3Value loss, falling not flatResponse length,not creeping upEntropy, not collapsing
Reward alone always looks good — it is reward rising while KL and length also climb that identifies hacking rather than learning.

The shape of the whole thing

Text
  preference pairs                    prompts only  (prompt, chosen, rejected)          (no responses needed)         |                                   |         v                                   |  [1] TRAIN REWARD MODEL                     |      base + scalar head                     |      Bradley-Terry pairwise loss            |      output: r(x, y) -> one float           |         |                                   |         +-------------> frozen -------------+                                             |                                             v                                 [2] PPO POLICY TRAINING                                     policy   (trainable, init = SFT)                                     value    (trainable)                                     ref      (frozen copy of SFT)                                     reward   (frozen, from step 1)                                             |                                             v                                       aligned model

Two stages, four models, and a strict rule: the reward model must be frozen during stage 2. If you train it concurrently on the policy's own outputs, the two models chase each other and the reward scale diverges within a few hundred steps.

Environment

Bash
pip install "trl>=1.0,<2" "transformers>=5" peft datasets \            accelerate bitsandbytes wandb scipy# The code in this lesson was checked against trl 1.13 and# transformers 5.17. Pin the versions you develop against and read# the docs for that exact version.pip freeze | grep -E "^(trl|transformers|peft)=="

That comment is not boilerplate. The reinforcement-learning trainers in this ecosystem have had breaking changes between releases more than once; TRL's PPO trainer was removed outright. Code copied from a blog post written a year earlier frequently fails with an import or argument error, so pin your versions and treat the installed docs as authoritative.

Stage 1: the reward model

Data preparation, with the traps closed

Python
from datasets import load_datasetfrom transformers import AutoTokenizerMODEL = "distilroberta-base"      # small enough to iterate quicklyMAXLEN = 512tok = AutoTokenizer.from_pretrained(MODEL)raw = load_dataset("Anthropic/hh-rlhf", split="train[:20000]")# hh-rlhf stores the whole conversation in "chosen" and "rejected";# RewardTrainer accepts these text columns as they are.# TRAP 1 + 2: never let one side be cut and the other not. Drop pairs# where EITHER side is too long, so truncation cannot become a# learnable signal. If you lose more than ~15%, raise MAXLEN instead.def fits(ex):    return (len(tok(ex["chosen"])["input_ids"])   <= MAXLEN and            len(tok(ex["rejected"])["input_ids"]) <= MAXLEN)ds = raw.filter(fits)print(f"dropped {100*(len(raw)-len(ds))/len(raw):.1f}% as too long")# TRAP 3: split by PROMPT, never by pair. Pairs sharing a prompt in# both train and eval let the model memorise good answers per prompt.ds = ds.train_test_split(test_size=0.05, seed=0)

Three traps, all of which silently inflate accuracy. Truncation asymmetry is the one from the opening. Prompt leakage is the subtler cousin: with 4 responses ranked per prompt you get 6 pairs, and a random split scatters them across train and test, so the model has already seen the prompt and most of its answers.

Current TRL RewardTrainer also drops, rather than truncates, pairs longer than max_length, but measure the loss yourself: if it removes a third of your data, you want to know. And train_test_split splits by row. With one pair per prompt, as in hh-rlhf, that is a split by prompt; with several pairs per prompt you must group by prompt first.

The model

Python
from transformers import AutoModelForSequenceClassificationrm = AutoModelForSequenceClassification.from_pretrained(    MODEL,    num_labels=1,               # ONE output: a scalar score    problem_type="regression",)# The head is randomly initialised. Keep it small so early rewards# start near zero rather than at some arbitrary large scale.rm.classifier.out_proj.weight.data.normal_(std=1e-3)rm.classifier.out_proj.bias.data.zero_()

num_labels=1 is what turns a classifier into a reward model. There is no softmax and no class; the forward pass returns a single unbounded real number per sequence. Anything above zero is not "good" and anything below is not "bad" — only differences between scores carry meaning, because the training objective sees only differences.

Why this loss, specifically

The training data is comparisons, not scores, so you need an objective that turns "A beat B" into pressure on two numbers. The Bradley-Terry model, borrowed from the statistics of paired contests, supplies it. It assumes the chance of preferring A is a sigmoid of the gap between their latent strengths:

P(yw≻yl∣x)=σ(rϕ(x,yw)−rϕ(x,yl))P(y_w \succ y_l \mid x) = \sigma\big(r_\phi(x, y_w) - r_\phi(x, y_l)\big)

Fitting by maximum likelihood gives the loss:

LRM=−E(x,yw,yl)∼D[log⁡σ(rϕ(x,yw)−rϕ(x,yl))]\mathcal{L}_{\text{RM}} = -\mathbb{E}_{(x, y_w, y_l) \sim \mathcal{D}} \Big[\log \sigma\big(r_\phi(x, y_w) - r_\phi(x, y_l)\big)\Big]

In plain terms: for every labelled pair, raise the winner's score above the loser's — and ease off once the gap is comfortable.

The shape earns its place for three reasons. It depends only on the difference, which is exactly the information a comparison carries and nothing more, so the model is never asked to invent an absolute scale that the data cannot support. It saturates, so once a pair is ranked correctly with a healthy gap the gradient nearly vanishes and capacity moves to pairs still getting it wrong. And it is smooth and convex in the gap, so optimisation is well behaved — unlike a hinge or a hard ranking constraint.

Work a batch of three pairs to see the gradient allocation:

Pairr(yw)r(y_w)r(yl)r(y_l)Gap ddσ(d)\sigma(d)Loss −log⁡σ(d)-\log\sigma(d)∂L/∂d=σ(d)−1\partial\mathcal{L}/\partial d = \sigma(d)-1
Confidently right2.100.301.800.8580.153-0.142
Barely right0.900.800.100.5250.644-0.475
Wrong-0.401.10-1.500.1821.701-0.818

Mean loss is (0.153+0.644+1.701)/3=0.833(0.153 + 0.644 + 1.701)/3 = 0.833. The gradient magnitudes tell the real story: the wrongly ranked pair pushes 5.8 times harder than the confidently correct one. The objective self-prioritises without you writing any curriculum logic.

Training it

Python
from trl import RewardTrainer, RewardConfigcfg = RewardConfig(    output_dir="rm-out",    per_device_train_batch_size=8,    gradient_accumulation_steps=4,     # effective batch 32    num_train_epochs=1,                # RMs overfit fast - start here    learning_rate=1e-5,    warmup_ratio=0.03,    lr_scheduler_type="cosine",    max_length=MAXLEN,    eval_strategy="steps",    eval_steps=100,    bf16=True,)trainer = RewardTrainer(    model=rm, args=cfg, processing_class=tok,    train_dataset=ds["train"],     # text columns: chosen, rejected    eval_dataset=ds["test"],)trainer.train()

One epoch is deliberate, and it has to be set explicitly: RewardConfig defaults to three epochs and a learning rate of 1e-4. On datasets below roughly 50,000 pairs, reward models begin memorising in the second epoch; the tell is training reward gaps widening while eval accuracy plateaus or slips. If you want the loss written out rather than delegated:

Python
import torch.nn.functional as Fdef pairwise_loss(model, batch):    r_w = model(input_ids=batch["input_ids_chosen"],                attention_mask=batch["attention_mask_chosen"]).logits.squeeze(-1)    r_l = model(input_ids=batch["input_ids_rejected"],                attention_mask=batch["attention_mask_rejected"]).logits.squeeze(-1)    # softplus(-d) == -log(sigmoid(d)), but does not underflow to -inf    loss = F.softplus(-(r_w - r_l)).mean()    return loss, (r_w > r_l).float().mean(), (r_w - r_l).mean()

Evaluating it — three checks, not one

Python
import numpy as np, torchfrom scipy.stats import pearsonr@torch.no_grad()def audit(model, tok, eval_ds, device="cuda"):    model.eval()    gaps, lens_w, lens_l, scores = [], [], [], []    for ex in eval_ds:        rw = score(model, tok, ex["chosen"], device)        rl = score(model, tok, ex["rejected"], device)        gaps.append(rw - rl)        lens_w.append(len(tok(ex["chosen"])["input_ids"]))        lens_l.append(len(tok(ex["rejected"])["input_ids"]))        scores += [rw, rl]    gaps = np.array(gaps)    lens = np.array(lens_w + lens_l)    acc      = float((gaps > 0).mean())    len_corr = float(pearsonr(scores, lens)[0])    # If you always picked the LONGER response, how often would you win?    len_base = float((np.array(lens_w) > np.array(lens_l)).mean())    print(f"pairwise accuracy   {acc:.3f}")    print(f"length baseline     {len_base:.3f}   <- beat this by a lot")    print(f"reward~length corr  {len_corr:+.3f}  <- keep below 0.20")    print(f"mean gap            {gaps.mean():+.3f}")    return acc, len_base, len_corr

The length baseline is the check that would have saved the team in the opening story, and almost nobody computes it. Suppose your reward model hits 0.71 accuracy. Impressive — unless a policy of "always pick the longer response" also scores 0.68 on the same data, in which case your model has contributed three points over a heuristic that requires no learning at all. Healthy numbers look like: accuracy 0.70, length baseline 0.55, length correlation 0.12.

SymptomLikely causeFix
Accuracy above 0.90Leakage, truncation artefact, or trivially easy pairsCheck the length baseline; verify split is by prompt; inspect 20 pairs by hand
Accuracy near 0.50Label noise, wrong pooling, or chosen/rejected swappedVerify labels on 30 examples; check the pooling index; print score pairs
Length correlation above 0.35Guidelines did not forbid length bias, or asymmetric generationFix at the data layer; length-match a subset and retrain
Eval accuracy falls after step NOverfittingStop at N; reduce to one epoch; lower the learning rate
Reward scale drifts to ±40Head initialised too large, or learning rate too highSmall head init; the scale is arbitrary but a wild one destabilises the value model later

Scoring text with it

Python
@torch.no_grad()def score(model, tok, text, device="cuda"):    enc = tok(text, return_tensors="pt", truncation=True,              max_length=MAXLEN).to(device)    return model(**enc).logits.squeeze().item()prompt = "Human: My laptop won't turn on. Assistant:"candidates = [    " Have you tried turning it off and on again?",    (" Let's isolate it. 1) Hold power for 15s to force a drain, then"     " plug in and retry. 2) Check the charger LED - no light means"     " charger or port. 3) Try a known-good charger. If the fan spins"     " but the screen stays dark, it's likely display or GPU."),    " I'm sorry to hear that. Computers can be frustrating sometimes!",]for c in candidates:    print(f"{score(rm, tok, prompt + c):+.3f}  {c[:52]}...")# Typical output:#   -0.412  Have you tried turning it off and on again?...#   +1.847  Let's isolate it. 1) Hold power for 15s to force...#   -0.938  I'm sorry to hear that. Computers can be frustra...

Do this on 20 hand-written candidates before you start PPO. It takes fifteen minutes and it is the only way to see whether the model has learned quality or a proxy. If the empathetic-but-useless response outscores the actionable one, your policy will learn to be empathetic and useless — with great efficiency.

Score twenty responses you wrote yourself and read the ordering. If the reward model disagrees with you there, every hour you then spend on PPO is spent making the policy better at being wrong.

Stage 2: what PPO is actually doing here

The reward model gives a score for a whole response. Gradient descent needs a per-token signal. PPO bridges that gap with three ideas.

Advantages instead of raw rewards. A value model predicts the expected return from each partial response, and the advantage is the difference between what happened and what was expected. This centres the signal: on an easy prompt where every response scores near 8.0, raw rewards say "all good" while advantages still separate the 8.4 from the 7.8.

A clipped objective so one batch can be reused. Generating rollouts is far more expensive than the gradient step, so you want several optimisation passes per rollout. But policy gradients are only valid on-policy. PPO corrects with an importance ratio and then bounds it:

LCLIP=Et[min⁡(ρtAt, clip(ρt, 1−ϵ, 1+ϵ) At)],ρt=πθ(at∣st)πθold(at∣st)L^{\text{CLIP}} = \mathbb{E}_t\Big[\min\big(\rho_t A_t,\ \text{clip}(\rho_t,\, 1-\epsilon,\, 1+\epsilon)\, A_t\big)\Big], \qquad \rho_t = \frac{\pi_\theta(a_t \mid s_t)}{\pi_{\theta_{\text{old}}}(a_t \mid s_t)}

Plainly: score the update both with the true importance ratio and with the ratio pinned inside [0.8,1.2][0.8, 1.2] , then keep the smaller. Taking the minimum removes any reward for pushing a token further once it has already moved a lot — while still allowing the full gradient to pull a token back if the policy moved it the wrong way.

A KL penalty folded into the reward. Clipping bounds each step but not cumulative drift. So every token pays a tax for how far the policy has moved from a frozen reference:

r~t=−βlog⁡πθ(at∣st)πref(at∣st)  +  1[t=T] rϕ(x,y)\tilde{r}_t = -\beta \log \frac{\pi_\theta(a_t \mid s_t)}{\pi_{\text{ref}}(a_t \mid s_t)} \;+\; \mathbb{1}[t = T]\, r_\phi(x, y)

Reading it: charge a small penalty at every token proportional to the local drift, and add the reward model's score only at the final token. Spreading the penalty across tokens rather than applying it in one lump at the end is what lets the value model attribute drift to the tokens that caused it.

Running PPO

Older tutorials run this stage with TRL's PPOTrainer and a ppo_trainer.step() loop. That code no longer runs: TRL 1.x no longer ships PPOTrainer or AutoModelForCausalLMWithValueHead, and its online trainers today are the critic-free GRPOTrainer and RLOOTrainer (section 3 covers them). For PPO with a critic and a learned reward model, current practice is a framework built around rollout generation: OpenRLHF or veRL. Both generate with vLLM, train with DeepSpeed or FSDP, and use Ray to place the four models on GPUs.

An OpenRLHF launch, trimmed to the flags that map onto this lesson. The flag names are from the project's README in 2026 and have been renamed between releases, so check them against the version you install.

Bash
pip install "openrlhf[vllm]"ray start --head --num-gpus 8ray job submit --address="http://127.0.0.1:8265" \   -- python3 -m openrlhf.cli.train_ppo_ray \   --actor.model_name_or_path ./sft-model \   --reward.model_name_or_path ./rm-out \   --data.prompt_dataset ./prompts \   --data.input_key prompt \   --data.apply_chat_template \   --rollout.batch_size 1024 \   --rollout.n_samples_per_prompt 1 \   --train.batch_size 128 \   --train.max_epochs 1 \   --generate_max_len 1024 \   --actor.adam.lr 5e-7 \   --critic.adam.lr 9e-6 \   --algo.kl.init_coef 0.01 \   --reward.normalize_enable \   --ds.zero_stage 3 \   --ds.param_dtype bf16

Map the flags back to the algorithm. --algo.kl.init_coef is β\beta. --reward.normalize_enable is the reward normalisation described below. The critic has its own, higher learning rate. Leaving --algo.advantage.estimator unset gives PPO with GAE; setting it to rloo or group_norm switches to the critic-free methods in section 3 without changing anything else. OpenRLHF expects a reward model in its own format, so train it with its openrlhf.cli.train_rm script, or pass --reward.remote_url pointing at a Python reward function instead.

veRL expresses the same run as python3 -m verl.trainer.main_ppo with Hydra-style overrides such as algorithm.adv_estimator=gae and algorithm.kl_ctrl.kl_coef=0.001. Its quickstart trains Qwen2.5-0.5B-Instruct on GSM8K on a single GPU, which makes it a good first PPO run.

Whichever framework you use, one iteration does the same thing. Written out:

Python
logp_old = logprobs_of(policy, queries, responses)   # AT GENERATION TIMElogp_ref = logprobs_of(ref,    queries, responses)   # frozen modelvalues   = value_head(policy, queries, responses)# per-token reward: KL tax everywhere, RM score at the last tokenkl        = logp_old - logp_refper_tok_r = -beta * klper_tok_r[:, -1] += rm_scores# GAEadv, returns = [], []last = 0.0for t in reversed(range(T)):    next_v = values[:, t + 1] if t < T - 1 else 0.0    delta  = per_tok_r[:, t] + gamma * next_v - values[:, t]    last   = delta + gamma * lam * last    adv.insert(0, last)adv     = torch.stack(adv, dim=1)returns = adv + valuesadv     = (adv - adv.mean()) / (adv.std() + 1e-8)   # per-batch whiteningfor _ in range(ppo_epochs):    for mb in minibatches(...):        logp_new = logprobs_of(policy, mb.queries, mb.responses)        rho      = torch.exp(logp_new - mb.logp_old)     # NOT recomputed        pg = -torch.min(rho * mb.adv,                        rho.clamp(1 - eps, 1 + eps) * mb.adv).mean()        v_clipped = mb.values + (value_head(...) - mb.values).clamp(-cv, cv)        vf = 0.5 * torch.max((value_head(...) - mb.returns) ** 2,                             (v_clipped - mb.returns) ** 2).mean()        loss = pg + vf_coef * vf        loss.backward()        torch.nn.utils.clip_grad_norm_(policy.parameters(), 1.0)        opt.step(); opt.zero_grad()

The comment on logp_old marks the most damaging bug in PPO implementations. Those log-probabilities must be the ones recorded when the tokens were generated. Recompute them from the current policy and every ratio is exactly 1.0 at the first inner epoch, clipping never engages, and you are running unconstrained policy gradient while your logs cheerfully report a clip fraction of zero.

Stability techniques and the failure each one prevents

TechniqueFailure it preventsHow
Reward normalisation (running mean/std)Value model divergenceThe RM's raw scale is arbitrary and may drift as the policy moves. A critic regressing on an unbounded moving target diverges; normalising to roughly zero mean, unit variance keeps it in range
Advantage whitening per batchWildly varying step sizesPuts advantages on a consistent scale so a batch of easy prompts does not produce a giant update
Adaptive KL coefficientKL explosion or total stagnationMeasures actual KL each step; if above target, raise β\beta; if below, lower it. Turns a brittle constant into a controller
Value clippingCritic overshootBounds how far the value prediction may move from the rollout-time value, mirroring the policy clip
Gradient norm clipping (1.0)Single-batch catastropheOne pathological batch cannot produce an update large enough to wreck the model
Sampling with varied lengths and temperature 1.0Learning stalls; length collapseWithout generation diversity all responses in a batch are near-identical, advantages are near-zero, and there is nothing to learn from
Reward-model ensemble, take the minimumReward hackingExploits usually fool one model, not three. The minimum is pessimistic precisely where the ensemble disagrees

Worked example of why reward normalisation matters. Suppose the RM outputs scores in the range [−3,+12][-3, +12] with mean 4.2 and standard deviation 3.1. The value head is randomly initialised and starts predicting near 0. Its first regression targets are around 4.2, giving early advantages of roughly +4+4 on essentially every token regardless of quality — so the first few hundred steps uniformly increase the probability of whatever the policy happened to sample. Normalise, and those same early advantages centre near zero, carrying only the relative information you actually want.

Every stability technique in a PPO implementation exists because someone shipped a run without it and watched a model destroy itself. None of them are optional decorations.

Monitoring: the six numbers to watch

MetricHealthyAlarmReading
Mean rewardRises steadily then plateausRises without limitUnbounded growth is hacking, not learning
KL to referenceSettles near your target, 4–10Past 30, or stuck below 1Too high: degenerate text. Too low: you trained nothing
Policy clip fraction0.05 – 0.25Above 0.4Too much clipping means steps too large or too many inner epochs
Value lossFalls then flattensRisingValue model diverging — normalise rewards, lower its LR
Mean response lengthRoughly stableClimbing steadilyThe single clearest fingerprint of length hacking
Policy entropyDeclines gentlyFalls off a cliffMode collapse; add an entropy bonus and raise β\beta

Each framework logs these under its own names (veRL, for example, uses actor/pg_clipfrac, critic/vf_loss and response_length/mean). Read them together, never alone. Reward up with length stable and KL near target is real progress. Reward up with length climbing and entropy falling is a model learning to exploit your scorer, and the reward curve will look identical in both cases.

Pitfalls worth naming

  • Reward model and policy on different tokenisers. Perfectly legal — they are separate models — but you must decode the policy's tokens to text and re-encode with the reward model's tokeniser. Feeding policy token IDs straight into the reward model produces scores from nonsense text, and nothing errors.
  • Special tokens leaking into the score. If the policy emits padding or an unexpected chat template token, the reward model has never seen that pattern and its score is arbitrary. Strip and normalise before scoring.
  • Greedy decoding during rollouts. Zero diversity, zero advantage spread, no learning. The training loop runs happily and the reward stays flat.
  • Forgetting to freeze the reference. If the reference model shares parameters with the policy, KL is identically zero, the leash is off, and you find out three hours in.
  • Training more than one epoch on a small reward-model dataset. The reward model memorises, and the policy then optimises against memorised quirks.
  • Judging success from the reward curve. Hold out prompts that never touch training and run a blind side-by-side against the SFT model every few hundred steps. That number, not the reward, decides whether you ship.

How to actually run this the first time

Debug the entire pipeline at toy scale before touching a GPU cluster. A 125M policy, a 125M reward model, 500 preference pairs, 50 PPO steps. Everything that will break structurally — tokeniser mismatch, reward shape errors, frozen-model mix-ups, KL identically zero — breaks here, in minutes, on one GPU. Scaling a working pipeline is routine; debugging a broken one at 7B is not.

Audit the reward model by hand before it trains anything. Twenty candidate responses you wrote yourself, scored and sorted. If the ordering matches your judgement, proceed. If it does not, no amount of PPO tuning will help, because the policy will faithfully optimise toward whatever the reward model actually believes.

Set an adaptive KL target and let it do the work. A target of 6 total nats per response is a sane default. Fixing β\beta by hand means retuning it every time the reward scale or response length changes; the controller absorbs that automatically.

Checkpoint every 100 steps and keep a manual sample log. Save five generations for the same fixed prompts at every checkpoint, in a plain text file you actually read. Degradation is obvious to a human two hundred steps before it is obvious in any metric, and the checkpoint you want is always the one just before you noticed.

Stop earlier than feels right. The best checkpoint by human evaluation is very often not the one with the highest reward. Training until the reward plateaus reliably overshoots into the region where the policy has begun exploiting the reward model rather than satisfying it.