Course Content
Reinforcement Learning from Human Feedback (RLHF)
4 sections · 10 lessons
Proximal Policy Optimization — The Clipped Objective
Here is a training run that fails. You have a language model, a trained reward model, and an obvious plan: sample a response, get its reward, and do gradient ascent to make high-reward responses more likely. You write the update, set the learning rate to 1e-5, and start.
For the first 200 steps the mean reward climbs from 0.4 to 1.1. Encouraging. At step 340 the reward jumps to 4.8. At step 500 it is 9.2 and rising. You sample the model to see what a 9.2-reward response looks like:
Prompt: How do I reset my router?Response: Certainly! Certainly! I'd be happy to help you with that, I'd be happy to help. Certainly! I'd be happy to help you with that, absolutely, certainly, I'd be happy to help...The reward model loves it. Every human hates it. Somewhere between step 200 and step 340 the policy stopped producing language and started producing an adversarial input to your reward function.
Proximal Policy Optimization is, at its heart, one specific answer to why this happened and how to stop it. Understanding PPO means understanding the failure first.
The language model as a reinforcement learning problem
Reinforcement learning has a standard vocabulary. Mapping it onto text generation is mostly mechanical, but two of the mappings are unusual enough to cause confusion later.
| RL term | Meaning | In language generation |
|---|---|---|
| State st | What the agent observes | The prompt plus all tokens generated so far |
| Action at | What the agent does | Emit the next token — one choice from a vocabulary of 32,000 to more than 150,000 tokens |
| Policy πθ(a∣s) | Probability of each action given the state | Exactly the model's softmax over the vocabulary. The language model is the policy |
| Episode | One run from start to terminal state | One complete response, from the first token to the end-of-sequence token |
| Reward rt | Scalar feedback | Zero for every token except the last, where the reward model's score for the whole response arrives |
Two things are unusual here. First, the action space is enormous — tens of thousands of discrete actions per step, compared to a handful in a typical control problem. Second, the reward is terminal and sparse: you generate 250 tokens and receive one number at the end. Nothing tells you which of those 250 token choices earned the score. This is the credit assignment problem, and it is why PPO for language models needs a value function.
Value and advantage
The value function V(st) answers: from this partial response, what total reward should I expect if I carry on as I currently do? It is a prediction, learned alongside the policy by regression against observed returns.
The advantage answers the more useful question: was this particular token better or worse than what I would normally have done here?
In words: the value of taking this specific action, minus the average value of being in this state at all. A positive advantage means the token beat expectations; a negative advantage means it fell short. Crucially, the advantage is a relative quantity, and that is what makes it a usable training signal.
Why relative? Suppose every response to a given prompt scores between 8.0 and 8.5 because the prompt is easy. Raw rewards say "everything is great" and give almost no signal about which response was better. Advantages centre this: the 8.5 response gets a positive advantage, the 8.0 response a negative one, and the policy learns the distinction. Subtracting the baseline V(st) dramatically reduces the variance of the gradient estimate without introducing bias.
Reward tells you how good the outcome was. Advantage tells you how much of that was your doing. Only the second is a training signal.
Generalised advantage estimation
In practice, advantages are computed with GAE, which trades bias against variance using a parameter λ:
Read it plainly: δt is the one-step surprise — how much better the outcome was than the value function predicted. GAE then sums these surprises forward through the episode, discounting each by γλ. With λ=0 you use only the immediate surprise (low variance, high bias, because you are trusting the value function completely). With λ=1 you use the full observed return (unbiased, high variance). RLHF typically uses λ=0.95 and γ=1.0 — no discounting, because a token at position 5 and a token at position 200 both contribute to the same single end-of-episode reward, and discounting would arbitrarily favour early tokens.
Why naive gradient ascent explodes
The simplest policy gradient, REINFORCE, is:
Meaning: increase the log-probability of actions with positive advantage, decrease it for negative advantage, in proportion to how large the advantage was. It is correct — it really is an unbiased estimate of the gradient of expected reward. It is also nearly unusable, for a reason that has nothing to do with correctness.
The problem is that this gradient is only valid for data collected under the current policy. The moment you take a step, the policy changes, and your collected samples were drawn from a distribution that no longer exists. Formally, policy gradient methods are on-policy: one batch of samples buys you exactly one gradient step.
That is catastrophically wasteful for language models. Generating a batch of 512 responses at 256 tokens each means producing 131,072 tokens in 256 sequential decoding steps, each a forward pass of a multi-billion-parameter model over the whole batch. On a serious cluster this takes minutes. Using all that generation for a single gradient step and then throwing it away means your GPUs spend the overwhelming majority of their time generating rather than learning.
The obvious fix — reuse the batch for several gradient steps — breaks the maths. After two or three updates the policy has moved, the samples are stale, and the gradient estimate is biased. And because the estimate is biased in a direction that looks like improvement, the optimiser happily marches off a cliff. That is the run described at the top: repeated updates on stale data, no constraint on step size, and a policy that drifts into a region where the reward model is nonsense.
The distribution shift, made concrete
Suppose your policy currently assigns probability 0.02 to a particular token, and one lucky sample gave it a large positive advantage of 8.0. An unconstrained gradient step could raise that probability to 0.4 — a twentyfold increase from a single observation. Do that across a few thousand tokens and the output distribution is unrecognisable. The reward model, which was trained on outputs from the original distribution, has no reliable opinion about the new text at all, and its confident-but-meaningless scores drive the next round of updates.
PPO's answer: importance sampling with a hard limit
PPO makes two moves. First, it makes off-policy reuse valid by importance sampling. Second, it prevents the reuse from going too far by clipping.
The importance ratio compares the current policy's probability for an action against the probability under the policy that actually generated the data:
A ratio of 1.0 means the policy has not changed for this token. A ratio of 1.3 means the current policy is 30% more likely to emit it than the sampling policy was. Weighting the objective by this ratio corrects for the mismatch — that is standard importance sampling.
The trouble is that importance sampling has exploding variance when the ratio is far from 1. A ratio of 20 multiplies that sample's contribution by 20, and one such sample can dominate an entire batch. So PPO clips:
Plain-English reading: compute the importance-weighted objective two ways — once normally, once with the ratio forcibly held inside [1−ϵ,1+ϵ] — and take whichever is smaller. Typically ϵ=0.2, so ratios are constrained to [0.8,1.2].
Taking the minimum is the essential trick, and it produces a deliberate asymmetry. Work the four cases with ϵ=0.2:
| Situation | ρt | At | ρA | clip(ρ)A | min | Gradient? |
|---|---|---|---|---|---|---|
| Good token, policy already boosted it a lot | 1.35 | +2.0 | 2.70 | 2.40 | 2.40 | No — clipped term wins, gradient is zero. Stop boosting. |
| Good token, policy has moved little | 1.05 | +2.0 | 2.10 | 2.10 | 2.10 | Yes — keep boosting |
| Bad token, policy already suppressed it a lot | 0.60 | -2.0 | -1.20 | -1.60 | -1.60 | No — clipped, gradient zero. Stop suppressing. |
| Bad token, but policy increased its probability | 1.35 | -2.0 | -2.70 | -2.40 | -2.70 | Yes — unclipped term wins, full gradient pushes it back down |
That last row is the point of the minimum. When the policy has drifted in the wrong direction — raising the probability of an action that turned out to be bad — clipping does not protect it. The unclipped term is more negative, the min selects it, and the full corrective gradient flows. Clipping caps how far you can be rewarded for a change, but never caps your ability to undo a mistake.
Clipping does not shrink the gradient. It switches it off entirely once a token has moved far enough — but only in the direction that would move it further.
The full PPO loss adds a value-function term and an entropy bonus:
The value loss trains the critic to predict returns accurately, because a bad critic produces bad advantages and therefore bad policy gradients. The entropy bonus, typically with c2 between 0.0 and 0.01, rewards keeping the output distribution spread out; it is the main defence against mode collapse, where the model narrows to one phrasing for everything.
The KL leash, and where it actually lives
Clipping bounds each update. It does nothing about cumulative drift: a thousand small legal steps in the same direction still arrive somewhere strange. So RLHF adds a second constraint measured against a frozen copy of the starting model — the reference model, which is the SFT checkpoint with its weights fixed.
The standard implementation folds the penalty into the per-token reward rather than the loss:
Reading it: every token pays a small tax proportional to how much more likely the current policy makes it than the frozen reference did, and only the final token also collects the reward model's score. Because the penalty is per-token, drift is charged continuously throughout the response rather than only at the end — which means the value function and advantages see it, and the credit assignment is far better than a single lump penalty would give.
Work a case. A response has 200 tokens, mean per-token log-ratio of 0.06, and reward model score 3.1, with β=0.1. Total KL is 200×0.06=12.0; the penalty is 0.1×12.0=1.2; net objective is 3.1−1.2=1.9. A more conservative response scores 2.4 with total KL 3.0, giving 2.4−0.3=2.1 — the conservative one wins. Now change β to 0.02: the first becomes 3.1−0.24=2.86, the second 2.4−0.06=2.34, and the aggressive response wins. Same model, same data, opposite training outcome from one hyperparameter.
Four models, one training loop
| Model | Purpose | Trained? | Memory |
|---|---|---|---|
| Policy πθ | Generates responses; the thing being aligned | Yes | Weights + gradients + Adam states ≈ 16 bytes/param |
| Reference πref | Frozen SFT copy; anchor for the KL penalty | No | Weights only, inference precision |
| Reward model rϕ | Scores completed responses | No | Weights only |
| Value model Vψ | Predicts returns; feeds advantage estimation | Yes | Weights + gradients + optimiser states |
For a 7B policy in bf16 with Adam, that is roughly 112 GB for the policy, 14 GB each for the frozen reference and reward model, and another 112 GB for the value model if it is the same size — before activations, before the KV cache for generation, before the rollout buffer. This is why practitioners share a backbone between policy and value with two heads, use LoRA adapters so the reference model is simply the base weights with adapters disabled, and run reward models an order of magnitude smaller than the policy.
The loop itself
repeat until done: # --- ROLLOUT (no gradients, this is generation) --- sample a batch of prompts generate responses with the CURRENT policy, temperature ~1.0 record log-probs under the policy at generation time -> pi_old score each response with the reward model -> r compute per-token log-probs under the reference -> pi_ref build per-token rewards: r_tilde = -beta*(log pi_old - log pi_ref) and add r at the final token run the value model over the sequences -> V compute advantages with GAE, then normalise them per batch # --- OPTIMISATION (gradients, reuse the same rollout) --- for epoch in 1..4: # typically 1-4 for minibatch in shuffle(rollout): rho = exp(logp_current - logp_old) policy_loss = -min(rho*A, clip(rho, 0.8, 1.2)*A).mean() value_loss = ((V_current - returns)**2).mean() loss = policy_loss + 0.1*value_loss - 0.01*entropy backprop, clip grad norm to 1.0, stepTwo details in that pseudo-code are easy to get wrong and expensive to debug. logp_old must be the log-probabilities captured at generation time, not recomputed later — recompute them after any update and every ratio becomes exactly 1.0, clipping never engages, and you have silently reverted to unconstrained gradient ascent. And advantage normalisation must be per batch, not global, or a batch of uniformly easy prompts will produce inflated advantages.
Hyperparameters that actually matter
| Parameter | Typical value | What happens if it is too high | Too low |
|---|---|---|---|
| Learning rate | 1e-6 to 5e-6 | KL explodes within a few hundred steps; gibberish | Nothing moves; reward flat for thousands of steps |
| KL coefficient β | 0.02 – 0.2 | Policy barely changes; reward flat — you have paid for RLHF and got the SFT model back | Reward soars, output quality collapses |
| Clip range ϵ | 0.2 | Larger steps, more instability | Very slow learning; most tokens clipped |
| PPO epochs per rollout | 1 – 4 | Data goes stale; ratios drift far from 1 and most gradient is clipped away | Wasted generation compute |
| Rollout batch size | 256 – 1024 prompts | Slow iteration | Noisy advantage estimates; unstable training |
| GAE λ | 0.95 | Higher variance advantages | Advantages inherit the value model's errors |
| Discount γ | 1.0 | — | Below 1.0 arbitrarily devalues later tokens for no principled reason |
Many implementations use an adaptive β: set a target KL (say 6.0 for the whole response), and if measured KL exceeds it, raise β; if it falls below, lower it. This turns a fragile hyperparameter into a controller with a setpoint, and is worth doing.
Four failure modes and their signatures
| Failure | What you observe | Cause | Fix |
|---|---|---|---|
| KL explosion | KL rises past 20–30 within a few hundred steps; outputs become repetitive or ungrammatical | β too small or learning rate too high; the leash is too long | Raise β, drop the learning rate, switch to adaptive KL with a target, clip gradient norm to 1.0 |
| Reward exploitation | Reward climbs steadily and smoothly; blind human evaluation gets worse; mean length or refusal rate moves sharply | The policy found a region where the reward model is wrong | Track surface statistics alongside reward; ensemble reward models and take the minimum; collect fresh preference data on current outputs |
| Value divergence | Value loss rises rather than falls; advantages become huge; policy loss oscillates | Value model chasing an unnormalised, drifting reward scale | Normalise rewards (running mean/std), clip value predictions to a range around the old value, use a lower LR for the value head |
| Mode collapse | Output entropy falls sharply; every response opens with the same phrase; distinct-n-gram counts drop | Over-optimisation toward a single reward peak | Add an entropy bonus, raise the KL penalty, stop earlier, and evaluate diversity as a first-class metric |
All four are detectable from the training logs, and none is detectable from the reward curve alone. A reward curve going up is consistent with every one of these.
Why PPO rather than the alternatives
| Method | How it constrains the step | Cost | Verdict for RLHF |
|---|---|---|---|
| REINFORCE | Not at all | Cheapest per step; one step per rollout | Unstable and sample-inefficient; workable only with heavy variance reduction and small steps |
| Vanilla actor-critic | Baseline reduces variance; no trust region | Moderate | Better than REINFORCE, still prone to destructive updates |
| TRPO | Hard KL constraint enforced via a second-order optimisation with conjugate gradients | Requires Fisher-vector products; painful with billions of parameters | Theoretically stronger guarantees, practically infeasible at LLM scale |
| PPO | Approximate trust region via first-order clipping | Standard backprop; a few lines of extra code | Good enough stability at ordinary cost — the reason it became the default for classic RLHF |
| RLOO / GRPO | PPO-style clipping (GRPO) or plain policy gradient (RLOO), with a baseline computed from several samples of the same prompt instead of a value model | No critic to train or hold in memory; needs several samples per prompt | The common choice since 2024, especially with verifiable rewards |
PPO is best understood as TRPO's guarantee, given up in exchange for something you can actually implement. TRPO enforces a genuine constraint on the KL between successive policies; PPO merely removes the incentive to violate it. That is weaker, and PPO can and does exceed the intended step size. But it needs only first-order gradients, and at seven billion parameters that difference decides which method exists in practice.
The value model has since come under the same scrutiny. Ahmadian et al. (2024) showed that for RLHF on language models, REINFORCE with a leave-one-out baseline (RLOO) can match or beat PPO while dropping the critic. GRPO, introduced in DeepSeekMath (Shao et al., 2024), keeps PPO's clipped objective but replaces the learned value with a group baseline. For each prompt it samples G responses, scores them, and gives every token of response i the same advantage:
In words: a response is good if it scored better than its siblings for the same prompt. With rewards of 1, 0, 0 and 1 from a correctness checker, the two correct answers get +1 and the two wrong ones −1 (using the population standard deviation, 0.5). There is no critic to diverge, which removes one of the failure modes listed above. The price is several generations per prompt, and no learning at all on prompts where every sample gets the same score. Later work (Dr. GRPO, DAPO) adjusted how the loss is averaged over tokens, because the original normalisation biased response length.
What this means when you run one
Instrument before you optimise. The minimum viable dashboard is: mean reward, KL from reference, mean response length, output entropy, value loss, and the fraction of tokens being clipped. Reward alone will tell you a lie you want to believe. If clip fraction is above roughly 0.3, your steps are too large or you are running too many epochs per rollout.
Start with a short leash and loosen it. Begin with β around 0.2 and a learning rate of 1e-6. A run that improves slowly can be accelerated; a run that has drifted into degenerate text cannot be recovered and you will restart from a checkpoint. The asymmetry of those two costs should determine your defaults.
Budget generation time, not just training time. In a typical PPO run, 60–80% of wall-clock time is spent generating rollouts, not computing gradients. Optimising the training step is usually the wrong place to look; batched generation with a fast inference engine, and a shorter maximum response length, usually buy far more.
Know when the answer is not PPO. The four-model memory footprint, the sensitivity to β and learning rate, and the failure modes above are real costs. If your reward is a checker or you can afford several samples per prompt, a critic-free method such as GRPO or RLOO removes the value model and its failure modes. If your preference dataset is fixed and moderate in size and you do not need on-policy sampling, a direct preference-optimisation method that removes the reward and value models entirely will get you most of the benefit for a fraction of the operational complexity. PPO earns its cost when you have a good reward model, need to keep sampling from the current policy, and want the control the KL leash gives you.