Reinforcement Learning from Human Feedback (RLHF)

Proximal Policy Optimization — The Clipped Objective


Here is a training run that fails. You have a language model, a trained reward model, and an obvious plan: sample a response, get its reward, and do gradient ascent to make high-reward responses more likely. You write the update, set the learning rate to 1e-5, and start.

For the first 200 steps the mean reward climbs from 0.4 to 1.1. Encouraging. At step 340 the reward jumps to 4.8. At step 500 it is 9.2 and rising. You sample the model to see what a 9.2-reward response looks like:

Text
Prompt:   How do I reset my router?Response: Certainly! Certainly! I'd be happy to help you with that,          I'd be happy to help. Certainly! I'd be happy to help you          with that, absolutely, certainly, I'd be happy to help...

The reward model loves it. Every human hates it. Somewhere between step 200 and step 340 the policy stopped producing language and started producing an adversarial input to your reward function.

Proximal Policy Optimization is, at its heart, one specific answer to why this happened and how to stop it. Understanding PPO means understanding the failure first.

The clip, on one token with advantage plus 20.550.850.951.001.101.35012345clippedto 0.8no change yetclippedto 1.2Cells hold the importance ratio: new policy probability over old, per token.
Outside the 0.8 to 1.2 band the objective goes flat, so a token the update wants to move hard contributes no gradient at all — the leash is in the loss, not the learning rate.

The language model as a reinforcement learning problem

Reinforcement learning has a standard vocabulary. Mapping it onto text generation is mostly mechanical, but two of the mappings are unusual enough to cause confusion later.

RL termMeaningIn language generation
State sts_tWhat the agent observesThe prompt plus all tokens generated so far
Action ata_tWhat the agent doesEmit the next token — one choice from a vocabulary of 32,000 to more than 150,000 tokens
Policy πθ(a∣s)\pi_\theta(a \mid s)Probability of each action given the stateExactly the model's softmax over the vocabulary. The language model is the policy
EpisodeOne run from start to terminal stateOne complete response, from the first token to the end-of-sequence token
Reward rtr_tScalar feedbackZero for every token except the last, where the reward model's score for the whole response arrives

Two things are unusual here. First, the action space is enormous — tens of thousands of discrete actions per step, compared to a handful in a typical control problem. Second, the reward is terminal and sparse: you generate 250 tokens and receive one number at the end. Nothing tells you which of those 250 token choices earned the score. This is the credit assignment problem, and it is why PPO for language models needs a value function.

Value and advantage

The value function V(st)V(s_t) answers: from this partial response, what total reward should I expect if I carry on as I currently do? It is a prediction, learned alongside the policy by regression against observed returns.

The advantage answers the more useful question: was this particular token better or worse than what I would normally have done here?

At=Q(st,at)−V(st)A_t = Q(s_t, a_t) - V(s_t)

In words: the value of taking this specific action, minus the average value of being in this state at all. A positive advantage means the token beat expectations; a negative advantage means it fell short. Crucially, the advantage is a relative quantity, and that is what makes it a usable training signal.

Why relative? Suppose every response to a given prompt scores between 8.0 and 8.5 because the prompt is easy. Raw rewards say "everything is great" and give almost no signal about which response was better. Advantages centre this: the 8.5 response gets a positive advantage, the 8.0 response a negative one, and the policy learns the distinction. Subtracting the baseline V(st)V(s_t) dramatically reduces the variance of the gradient estimate without introducing bias.

Reward tells you how good the outcome was. Advantage tells you how much of that was your doing. Only the second is a training signal.

Generalised advantage estimation

In practice, advantages are computed with GAE, which trades bias against variance using a parameter λ\lambda:

δt=rt+γV(st+1)−V(st),AtGAE=∑k=0T−t−1(γλ)k δt+k\delta_t = r_t + \gamma V(s_{t+1}) - V(s_t), \qquad A_t^{\text{GAE}} = \sum_{k=0}^{T-t-1} (\gamma\lambda)^k \, \delta_{t+k}

Read it plainly: δt\delta_t is the one-step surprise — how much better the outcome was than the value function predicted. GAE then sums these surprises forward through the episode, discounting each by γλ\gamma\lambda. With λ=0\lambda = 0 you use only the immediate surprise (low variance, high bias, because you are trusting the value function completely). With λ=1\lambda = 1 you use the full observed return (unbiased, high variance). RLHF typically uses λ=0.95\lambda = 0.95 and γ=1.0\gamma = 1.0 — no discounting, because a token at position 5 and a token at position 200 both contribute to the same single end-of-episode reward, and discounting would arbitrarily favour early tokens.

Why naive gradient ascent explodes

The simplest policy gradient, REINFORCE, is:

∇θJ(θ)=E[∇θlog⁡πθ(at∣st) At]\nabla_\theta J(\theta) = \mathbb{E}\big[\nabla_\theta \log \pi_\theta(a_t \mid s_t) \, A_t\big]

Meaning: increase the log-probability of actions with positive advantage, decrease it for negative advantage, in proportion to how large the advantage was. It is correct — it really is an unbiased estimate of the gradient of expected reward. It is also nearly unusable, for a reason that has nothing to do with correctness.

The problem is that this gradient is only valid for data collected under the current policy. The moment you take a step, the policy changes, and your collected samples were drawn from a distribution that no longer exists. Formally, policy gradient methods are on-policy: one batch of samples buys you exactly one gradient step.

That is catastrophically wasteful for language models. Generating a batch of 512 responses at 256 tokens each means producing 131,072 tokens in 256 sequential decoding steps, each a forward pass of a multi-billion-parameter model over the whole batch. On a serious cluster this takes minutes. Using all that generation for a single gradient step and then throwing it away means your GPUs spend the overwhelming majority of their time generating rather than learning.

The obvious fix — reuse the batch for several gradient steps — breaks the maths. After two or three updates the policy has moved, the samples are stale, and the gradient estimate is biased. And because the estimate is biased in a direction that looks like improvement, the optimiser happily marches off a cliff. That is the run described at the top: repeated updates on stale data, no constraint on step size, and a policy that drifts into a region where the reward model is nonsense.

The distribution shift, made concrete

Suppose your policy currently assigns probability 0.02 to a particular token, and one lucky sample gave it a large positive advantage of 8.0. An unconstrained gradient step could raise that probability to 0.4 — a twentyfold increase from a single observation. Do that across a few thousand tokens and the output distribution is unrecognisable. The reward model, which was trained on outputs from the original distribution, has no reliable opinion about the new text at all, and its confident-but-meaningless scores drive the next round of updates.

PPO's answer: importance sampling with a hard limit

PPO makes two moves. First, it makes off-policy reuse valid by importance sampling. Second, it prevents the reuse from going too far by clipping.

The importance ratio compares the current policy's probability for an action against the probability under the policy that actually generated the data:

ρt(θ)=πθ(at∣st)πθold(at∣st)\rho_t(\theta) = \frac{\pi_\theta(a_t \mid s_t)}{\pi_{\theta_{\text{old}}}(a_t \mid s_t)}

A ratio of 1.0 means the policy has not changed for this token. A ratio of 1.3 means the current policy is 30% more likely to emit it than the sampling policy was. Weighting the objective by this ratio corrects for the mismatch — that is standard importance sampling.

The trouble is that importance sampling has exploding variance when the ratio is far from 1. A ratio of 20 multiplies that sample's contribution by 20, and one such sample can dominate an entire batch. So PPO clips:

LCLIP(θ)=Et[min⁡(ρt(θ) At,  clip(ρt(θ), 1−ϵ, 1+ϵ) At)]L^{\text{CLIP}}(\theta) = \mathbb{E}_t\Big[\min\big(\rho_t(\theta)\, A_t,\; \text{clip}(\rho_t(\theta),\, 1-\epsilon,\, 1+\epsilon)\, A_t\big)\Big]

Plain-English reading: compute the importance-weighted objective two ways — once normally, once with the ratio forcibly held inside [1−ϵ,1+ϵ][1-\epsilon, 1+\epsilon] — and take whichever is smaller. Typically ϵ=0.2\epsilon = 0.2, so ratios are constrained to [0.8,1.2][0.8, 1.2].

Taking the minimum is the essential trick, and it produces a deliberate asymmetry. Work the four cases with ϵ=0.2\epsilon = 0.2:

Situationρt\rho_tAtA_tρA\rho Aclip(ρ)A\text{clip}(\rho)AminGradient?
Good token, policy already boosted it a lot1.35+2.02.702.402.40No — clipped term wins, gradient is zero. Stop boosting.
Good token, policy has moved little1.05+2.02.102.102.10Yes — keep boosting
Bad token, policy already suppressed it a lot0.60-2.0-1.20-1.60-1.60No — clipped, gradient zero. Stop suppressing.
Bad token, but policy increased its probability1.35-2.0-2.70-2.40-2.70Yes — unclipped term wins, full gradient pushes it back down

That last row is the point of the minimum. When the policy has drifted in the wrong direction — raising the probability of an action that turned out to be bad — clipping does not protect it. The unclipped term is more negative, the min selects it, and the full corrective gradient flows. Clipping caps how far you can be rewarded for a change, but never caps your ability to undo a mistake.

Clipping does not shrink the gradient. It switches it off entirely once a token has moved far enough — but only in the direction that would move it further.

The full PPO loss adds a value-function term and an entropy bonus:

L(θ,ψ)=−LCLIP(θ)+c1Et[(Vψ(st)−Rt)2]⏟value loss−c2Et[H[πθ(⋅∣st)]]⏟entropy bonusL(\theta,\psi) = -L^{\text{CLIP}}(\theta) + c_1 \underbrace{\mathbb{E}_t\big[(V_\psi(s_t) - R_t)^2\big]}_{\text{value loss}} - c_2 \underbrace{\mathbb{E}_t\big[H[\pi_\theta(\cdot \mid s_t)]\big]}_{\text{entropy bonus}}

The value loss trains the critic to predict returns accurately, because a bad critic produces bad advantages and therefore bad policy gradients. The entropy bonus, typically with c2c_2 between 0.0 and 0.01, rewards keeping the output distribution spread out; it is the main defence against mode collapse, where the model narrows to one phrasing for everything.

The KL leash, and where it actually lives

Clipping bounds each update. It does nothing about cumulative drift: a thousand small legal steps in the same direction still arrive somewhere strange. So RLHF adds a second constraint measured against a frozen copy of the starting model — the reference model, which is the SFT checkpoint with its weights fixed.

The standard implementation folds the penalty into the per-token reward rather than the loss:

r~t={−βlog⁡πθ(at∣st)πref(at∣st)t<Trϕ(x,y)−βlog⁡πθ(aT∣sT)πref(aT∣sT)t=T\tilde{r}_t = \begin{cases} -\beta \log \dfrac{\pi_\theta(a_t \mid s_t)}{\pi_{\text{ref}}(a_t \mid s_t)} & t \lt T \\[2mm] r_\phi(x, y) - \beta \log \dfrac{\pi_\theta(a_T \mid s_T)}{\pi_{\text{ref}}(a_T \mid s_T)} & t = T \end{cases}

Reading it: every token pays a small tax proportional to how much more likely the current policy makes it than the frozen reference did, and only the final token also collects the reward model's score. Because the penalty is per-token, drift is charged continuously throughout the response rather than only at the end — which means the value function and advantages see it, and the credit assignment is far better than a single lump penalty would give.

Work a case. A response has 200 tokens, mean per-token log-ratio of 0.06, and reward model score 3.1, with β=0.1\beta = 0.1. Total KL is 200×0.06=12.0200 \times 0.06 = 12.0; the penalty is 0.1×12.0=1.20.1 \times 12.0 = 1.2; net objective is 3.1−1.2=1.93.1 - 1.2 = 1.9. A more conservative response scores 2.4 with total KL 3.0, giving 2.4−0.3=2.12.4 - 0.3 = 2.1 — the conservative one wins. Now change β\beta to 0.02: the first becomes 3.1−0.24=2.863.1 - 0.24 = 2.86, the second 2.4−0.06=2.342.4 - 0.06 = 2.34, and the aggressive response wins. Same model, same data, opposite training outcome from one hyperparameter.

Four models, one training loop

ModelPurposeTrained?Memory
Policy πθ\pi_\thetaGenerates responses; the thing being alignedYesWeights + gradients + Adam states ≈ 16 bytes/param
Reference πref\pi_{\text{ref}}Frozen SFT copy; anchor for the KL penaltyNoWeights only, inference precision
Reward model rϕr_\phiScores completed responsesNoWeights only
Value model VψV_\psiPredicts returns; feeds advantage estimationYesWeights + gradients + optimiser states

For a 7B policy in bf16 with Adam, that is roughly 112 GB for the policy, 14 GB each for the frozen reference and reward model, and another 112 GB for the value model if it is the same size — before activations, before the KV cache for generation, before the rollout buffer. This is why practitioners share a backbone between policy and value with two heads, use LoRA adapters so the reference model is simply the base weights with adapters disabled, and run reward models an order of magnitude smaller than the policy.

The loop itself

Text
repeat until done:  # --- ROLLOUT (no gradients, this is generation) ---  sample a batch of prompts  generate responses with the CURRENT policy, temperature ~1.0  record log-probs under the policy at generation time -> pi_old  score each response with the reward model      -> r  compute per-token log-probs under the reference -> pi_ref  build per-token rewards: r_tilde = -beta*(log pi_old - log pi_ref)                           and add r at the final token  run the value model over the sequences         -> V  compute advantages with GAE, then normalise them per batch  # --- OPTIMISATION (gradients, reuse the same rollout) ---  for epoch in 1..4:                 # typically 1-4    for minibatch in shuffle(rollout):      rho = exp(logp_current - logp_old)      policy_loss = -min(rho*A, clip(rho, 0.8, 1.2)*A).mean()      value_loss  = ((V_current - returns)**2).mean()      loss = policy_loss + 0.1*value_loss - 0.01*entropy      backprop, clip grad norm to 1.0, step

Two details in that pseudo-code are easy to get wrong and expensive to debug. logp_old must be the log-probabilities captured at generation time, not recomputed later — recompute them after any update and every ratio becomes exactly 1.0, clipping never engages, and you have silently reverted to unconstrained gradient ascent. And advantage normalisation must be per batch, not global, or a batch of uniformly easy prompts will produce inflated advantages.

Hyperparameters that actually matter

ParameterTypical valueWhat happens if it is too highToo low
Learning rate1e-6 to 5e-6KL explodes within a few hundred steps; gibberishNothing moves; reward flat for thousands of steps
KL coefficient β\beta0.02 – 0.2Policy barely changes; reward flat — you have paid for RLHF and got the SFT model backReward soars, output quality collapses
Clip range ϵ\epsilon0.2Larger steps, more instabilityVery slow learning; most tokens clipped
PPO epochs per rollout1 – 4Data goes stale; ratios drift far from 1 and most gradient is clipped awayWasted generation compute
Rollout batch size256 – 1024 promptsSlow iterationNoisy advantage estimates; unstable training
GAE λ\lambda0.95Higher variance advantagesAdvantages inherit the value model's errors
Discount γ\gamma1.0—Below 1.0 arbitrarily devalues later tokens for no principled reason

Many implementations use an adaptive β\beta: set a target KL (say 6.0 for the whole response), and if measured KL exceeds it, raise β\beta; if it falls below, lower it. This turns a fragile hyperparameter into a controller with a setpoint, and is worth doing.

Four failure modes and their signatures

FailureWhat you observeCauseFix
KL explosionKL rises past 20–30 within a few hundred steps; outputs become repetitive or ungrammaticalβ\beta too small or learning rate too high; the leash is too longRaise β\beta, drop the learning rate, switch to adaptive KL with a target, clip gradient norm to 1.0
Reward exploitationReward climbs steadily and smoothly; blind human evaluation gets worse; mean length or refusal rate moves sharplyThe policy found a region where the reward model is wrongTrack surface statistics alongside reward; ensemble reward models and take the minimum; collect fresh preference data on current outputs
Value divergenceValue loss rises rather than falls; advantages become huge; policy loss oscillatesValue model chasing an unnormalised, drifting reward scaleNormalise rewards (running mean/std), clip value predictions to a range around the old value, use a lower LR for the value head
Mode collapseOutput entropy falls sharply; every response opens with the same phrase; distinct-n-gram counts dropOver-optimisation toward a single reward peakAdd an entropy bonus, raise the KL penalty, stop earlier, and evaluate diversity as a first-class metric

All four are detectable from the training logs, and none is detectable from the reward curve alone. A reward curve going up is consistent with every one of these.

Why PPO rather than the alternatives

MethodHow it constrains the stepCostVerdict for RLHF
REINFORCENot at allCheapest per step; one step per rolloutUnstable and sample-inefficient; workable only with heavy variance reduction and small steps
Vanilla actor-criticBaseline reduces variance; no trust regionModerateBetter than REINFORCE, still prone to destructive updates
TRPOHard KL constraint enforced via a second-order optimisation with conjugate gradientsRequires Fisher-vector products; painful with billions of parametersTheoretically stronger guarantees, practically infeasible at LLM scale
PPOApproximate trust region via first-order clippingStandard backprop; a few lines of extra codeGood enough stability at ordinary cost — the reason it became the default for classic RLHF
RLOO / GRPOPPO-style clipping (GRPO) or plain policy gradient (RLOO), with a baseline computed from several samples of the same prompt instead of a value modelNo critic to train or hold in memory; needs several samples per promptThe common choice since 2024, especially with verifiable rewards

PPO is best understood as TRPO's guarantee, given up in exchange for something you can actually implement. TRPO enforces a genuine constraint on the KL between successive policies; PPO merely removes the incentive to violate it. That is weaker, and PPO can and does exceed the intended step size. But it needs only first-order gradients, and at seven billion parameters that difference decides which method exists in practice.

The value model has since come under the same scrutiny. Ahmadian et al. (2024) showed that for RLHF on language models, REINFORCE with a leave-one-out baseline (RLOO) can match or beat PPO while dropping the critic. GRPO, introduced in DeepSeekMath (Shao et al., 2024), keeps PPO's clipped objective but replaces the learned value with a group baseline. For each prompt it samples GG responses, scores them, and gives every token of response ii the same advantage:

A^i=ri−mean⁡(r1,…,rG)std⁡(r1,…,rG)\hat{A}_i = \frac{r_i - \operatorname{mean}(r_1, \dots, r_G)}{\operatorname{std}(r_1, \dots, r_G)}

In words: a response is good if it scored better than its siblings for the same prompt. With rewards of 1, 0, 0 and 1 from a correctness checker, the two correct answers get +1+1 and the two wrong ones −1-1 (using the population standard deviation, 0.5). There is no critic to diverge, which removes one of the failure modes listed above. The price is several generations per prompt, and no learning at all on prompts where every sample gets the same score. Later work (Dr. GRPO, DAPO) adjusted how the loss is averaged over tokens, because the original normalisation biased response length.

What this means when you run one

Instrument before you optimise. The minimum viable dashboard is: mean reward, KL from reference, mean response length, output entropy, value loss, and the fraction of tokens being clipped. Reward alone will tell you a lie you want to believe. If clip fraction is above roughly 0.3, your steps are too large or you are running too many epochs per rollout.

Start with a short leash and loosen it. Begin with β\beta around 0.2 and a learning rate of 1e-6. A run that improves slowly can be accelerated; a run that has drifted into degenerate text cannot be recovered and you will restart from a checkpoint. The asymmetry of those two costs should determine your defaults.

Budget generation time, not just training time. In a typical PPO run, 60–80% of wall-clock time is spent generating rollouts, not computing gradients. Optimising the training step is usually the wrong place to look; batched generation with a fast inference engine, and a shorter maximum response length, usually buy far more.

Know when the answer is not PPO. The four-model memory footprint, the sensitivity to β\beta and learning rate, and the failure modes above are real costs. If your reward is a checker or you can afford several samples per prompt, a critic-free method such as GRPO or RLOO removes the value model and its failure modes. If your preference dataset is fixed and moderate in size and you do not need on-policy sampling, a direct preference-optimisation method that removes the reward and value models entirely will get you most of the benefit for a fraction of the operational complexity. PPO earns its cost when you have a good reward model, need to keep sampling from the current policy, and want the control the KL leash gives you.