Course Content
Reinforcement Learning from Human Feedback (RLHF)
4 sections · 10 lessons
Direct Preference Optimization (DPO)
Count the machinery a full PPO alignment run needs. A policy being trained. A frozen reference model. A reward model. A value model. Four sets of weights resident at once, generation inside the training loop, and a KL coefficient that will silently ruin the run if it is off by a factor of three. A single 7B pipeline can want 250 GB of accelerator memory before you have generated a token.
Now look at what that machinery ultimately produces. The reward model exists to turn preference pairs into a scalar. The scalar exists to guide the policy. The KL term exists to stop the policy leaving the region where the scalar is trustworthy. At the end you have a policy — and you throw the reward model away.
Direct Preference Optimization begins with a question: if the reward model is a disposable intermediate, is there a way to skip building it? The answer turns out to be yes, and not by approximation. Under the same assumptions PPO already makes, the language model is the reward model, expressed differently — and once you see that, the preference pairs can train the policy directly with an ordinary supervised loss.
| PPO pipeline | DPO pipeline | |
|---|---|---|
| Models in memory | Policy, reference, reward, value | Policy, reference |
| Generation during training | Yes — dominates wall-clock time | No |
| Training loop | Rollout, score, GAE, clipped update | One forward pass, one loss, backprop |
| Key hyperparameters | β, learning rate, clip range, GAE λ, epochs per rollout, value coefficient | β, learning rate |
| Typical time to a first working run | Days | Hours |
The derivation, one step at a time
Five steps. Each is short, and the payoff at step 4 is worth the effort of following steps 1 to 3.
Step 1 — write down what RLHF is optimising
The objective every KL-constrained RLHF method maximises is:
In words: get as much reward as possible, but pay a penalty proportional to how far the output distribution has moved from the frozen starting model. The coefficient β prices that movement. PPO approximates a solution to this by sampling and taking constrained gradient steps.
Step 2 — this objective has an exact solution
You do not have to approximate it. Treating it as a constrained optimisation over distributions and solving gives a closed form:
Reading it: the optimal policy is the reference model, reweighted by the exponential of reward. Responses the reference already found plausible and that score well get boosted; responses that score badly get suppressed; β controls how sharply. With β large the exponent is small and π∗ barely differs from πref; with β near zero it collapses onto whatever maximises reward.
This formula is exact but useless as an algorithm, because Z(x) sums over every possible response — an astronomically large set. That intractability is precisely why PPO exists.
Step 3 — turn it inside out
Here is the move. That equation relates a policy to a reward. Solve it for the reward instead. Take logs of both sides and rearrange:
This says something surprising: any policy implicitly defines a reward function. The reward a policy "believes in" is β times the log-ratio of its own probability to the reference's, plus a term that depends only on the prompt.
Z(x) is still intractable. But notice it depends only on x — the same value for every response to that prompt.
Step 4 — the cancellation
Preferences are modelled with Bradley-Terry, which says the probability a human prefers yw to yl is a sigmoid of the reward difference:
Substitute the expression from step 3 for both rewards:
The βlogZ(x) terms are identical and cancel exactly. Because Bradley-Terry only ever sees differences, the intractable partition function vanishes — not approximated, not bounded, simply gone. The preference probability becomes:
The reward model was never fundamentally necessary. It was a way of parameterising something the policy and the reference already encode between them, and the only quantity that made it intractable cancels the moment you take a difference.
Step 5 — read it as a loss
Replace π∗ with the policy you are training, πθ, and fit by maximum likelihood on your preference dataset:
Everything in it is computable with two forward passes. There is no sampling, no reward model, no value function, no rollouts. It is a classification loss over pairs.
Reading the loss without the notation
Define the implicit reward of a response as r^(x,y)=βlogπref(y∣x)πθ(y∣x) — how much more likely the policy makes this response than the reference did, scaled by β. The loss is then simply −logσ(r^w−r^l): make the policy relatively more enthusiastic about the winner than about the loser, compared with where it started.
The gradient makes the mechanism concrete:
Three things to notice. The bracketed part is ordinary maximum-likelihood on the winner minus maximum-likelihood on the loser — the same gradients supervised fine-tuning uses, with one sign flipped. The scalar in front is an automatic difficulty weighting: it is large when the model currently gets the pair backwards and small when the pair is already comfortably ranked. And the reference model appears only inside that weight, which is what keeps the policy tethered without any separate KL term.
A fully worked example
Take β=0.1 and one training pair. Sequence log-probabilities are sums over response tokens, so they are large negative numbers.
| Quantity | Chosen yw | Rejected yl |
|---|---|---|
| logπθ(y∣x) | -32.0 | -40.0 |
| logπref(y∣x) | -35.0 | -38.0 |
| Log-ratio | −32.0−(−35.0)=+3.0 | −40.0−(−38.0)=−2.0 |
| Implicit reward r^=β× log-ratio | 0.1×3.0=+0.30 | 0.1×(−2.0)=−0.20 |
Margin: r^w−r^l=0.30−(−0.20)=0.50. Loss: −logσ(0.50)=−log(0.6225)=0.474. Gradient weight: σ(−0.50)=0.3775.
Now the same arithmetic in three other regimes:
| Situation | Margin | Loss | Gradient weight σ(−margin) | Interpretation |
|---|---|---|---|---|
| Step 0 — policy is an exact copy of the reference | 0.00 | 0.693 | 0.500 | Both log-ratios are exactly zero. Every DPO run starts at loss log2=0.693 |
| Learning in progress | 0.50 | 0.474 | 0.378 | Correct but not confident; substantial gradient |
| Pair mastered | 3.00 | 0.049 | 0.047 | Gradient is 8x smaller than the previous row — the loss has stopped caring |
| Pair ranked backwards | -1.50 | 1.701 | 0.818 | Largest gradient in the batch; the update flips the ordering |
That first row is the single most useful diagnostic in DPO. Your loss must start at 0.6931. If it does not, the policy and reference are not identical at initialisation — you loaded different checkpoints, or the reference is not frozen, or the log-probabilities are being computed over different token spans.
Why β changes everything
Recompute the first example with different β, holding the log-ratios at +3.0 and −2.0:
| β | r^w | r^l | Margin | Loss | Behaviour |
|---|---|---|---|---|---|
| 0.01 | 0.03 | -0.02 | 0.05 | 0.668 | Loss barely responds to a large probability shift — the policy can drift enormously before the loss objects. Degenerate outputs |
| 0.1 | 0.30 | -0.20 | 0.50 | 0.474 | The standard default |
| 0.5 | 1.50 | -1.00 | 2.50 | 0.079 | Loss is already nearly satisfied; the policy stays very close to the reference and learns little |
β in DPO plays exactly the role the KL coefficient plays in PPO: it is the price of moving away from the reference. Small β means movement is cheap and the model drifts; large β means movement is expensive and the model does not change. Sensible values sit between 0.05 and 0.5, with 0.1 the usual starting point.
Implementing it correctly
1import torch2import torch.nn.functional as F34def sequence_logprob(model, input_ids, attention_mask, prompt_lens):5 """Sum of log-probs over RESPONSE tokens only."""6 logits = model(input_ids=input_ids,7 attention_mask=attention_mask).logits8 # Shift: logits at position t predict the token at position t+1.9 logits = logits[:, :-1, :]10 labels = input_ids[:, 1:].clone()1112 # Mask out prompt tokens AND padding. Both are essential.13 mask = attention_mask[:, 1:].clone().bool()14 for i, plen in enumerate(prompt_lens):15 mask[i, :max(plen - 1, 0)] = False1617 per_tok = torch.gather(18 logits.log_softmax(-1), dim=2, index=labels.unsqueeze(2)19 ).squeeze(2)20 return (per_tok * mask).sum(-1) # SUM, not mean212223def dpo_loss(policy, ref, batch, beta=0.1):24 pol_w = sequence_logprob(policy, batch["w_ids"], batch["w_mask"],25 batch["prompt_lens"])26 pol_l = sequence_logprob(policy, batch["l_ids"], batch["l_mask"],27 batch["prompt_lens"])28 with torch.no_grad(): # reference never trains29 ref_w = sequence_logprob(ref, batch["w_ids"], batch["w_mask"],30 batch["prompt_lens"])31 ref_l = sequence_logprob(ref, batch["l_ids"], batch["l_mask"],32 batch["prompt_lens"])3334 r_w = beta * (pol_w - ref_w) # implicit rewards35 r_l = beta * (pol_l - ref_l)36 loss = -F.logsigmoid(r_w - r_l).mean()3738 return loss, {39 "reward_accuracy": (r_w > r_l).float().mean().item(),40 "reward_margin": (r_w - r_l).mean().item(),41 "reward_chosen": r_w.mean().item(),42 "reward_rejected": r_l.mean().item(),43 }Four lines in that function are where implementations go wrong.
| Mistake | What happens | How you spot it |
|---|---|---|
| Not masking prompt tokens | The shared prompt's log-probs are added to both sides. They mostly cancel in the difference, but the reference-model terms do not cancel cleanly and gradients flow into predicting the prompt | Chosen and rejected implicit rewards both drift strongly in the same direction |
| Averaging per-token instead of summing | You have changed the objective. It is no longer the DPO loss, and length behaviour changes completely — this is a different algorithm (a reasonable one, but not this one) | Loss does not start at 0.693 under a longer/shorter pair; results differ from published baselines |
Reference model not in torch.no_grad() or not in eval() | Gradients flow into the reference, or dropout makes reference log-probs stochastic, adding noise to every margin | Loss is noisy at step 0 instead of exactly 0.693 |
| Off-by-one in the logits/labels shift | Every log-probability is for the wrong token | Reward accuracy hovers at 0.5 and never moves |
The summing decision deserves elaboration because it drives DPO's best-known weakness. A 300-token response has a sequence log-probability roughly three times more negative than a 100-token one. Small per-token probability changes therefore produce much larger log-ratio swings on long responses. If your chosen responses are systematically longer than your rejected ones, DPO will find that increasing length is a cheap way to raise the margin — and you get the same length inflation that plagues reward models, arriving through a different door. Length-match your pairs where you can, and always report mean chosen and rejected lengths in your dataset statistics.
Preparing the data
DPO wants three columns — prompt, chosen, rejected — but the details of how you build them decide whether the run works.
1def build_pair(example, tok):2 # 1. Format the prompt with the SAME chat template the SFT stage3 # used. A mismatch here is the most common cause of a run that4 # trains cleanly and produces a worse model: the policy is5 # being scored on a prompt format it was never trained on.6 prompt = tok.apply_chat_template(7 [{"role": "user", "content": example["question"]}],8 tokenize=False, add_generation_prompt=True)910 chosen = example["chosen"] + tok.eos_token11 rejected = example["rejected"] + tok.eos_token1213 # 2. Tokenise prompt and response SEPARATELY, then concatenate,14 # so you know exactly where the prompt ends. Tokenising the15 # joined string can merge the boundary tokens and shift the16 # mask by one, which silently corrupts every log-probability.17 p_ids = tok(prompt, add_special_tokens=False)["input_ids"]18 c_ids = tok(chosen, add_special_tokens=False)["input_ids"]19 r_ids = tok(rejected, add_special_tokens=False)["input_ids"]2021 # 3. Truncate the PROMPT from the left, never the response.22 # Cutting a response changes what you are comparing.23 if len(p_ids) > MAX_PROMPT:24 p_ids = p_ids[-MAX_PROMPT:]2526 return {"w_ids": p_ids + c_ids, "l_ids": p_ids + r_ids,27 "prompt_len": len(p_ids)}Three sanity checks before training, each of which takes a minute and each of which has saved runs: decode one w_ids and one l_ids back to text and read them, confirming the prompt appears exactly once and the response follows cleanly; assert that prompt_len is identical for the chosen and rejected members of every pair; and report the mean chosen and rejected response lengths across the dataset.
One more decision matters. DPO assumes the policy already assigns reasonable probability to responses like the chosen ones. If you run it directly on a base model that has never been instruction-tuned, both log-probabilities are tiny, the margins are dominated by noise, and training is unstable. The standard practice is to fine-tune on the chosen responses first — an ordinary SFT pass — and use that checkpoint as both the starting policy and the frozen reference.
Using TRL
1import torch2from trl import DPOTrainer, DPOConfig3from transformers import AutoModelForCausalLM, AutoTokenizer4from peft import LoraConfig56tok = AutoTokenizer.from_pretrained("sft-model")7policy = AutoModelForCausalLM.from_pretrained("sft-model",8 dtype=torch.bfloat16)910cfg = DPOConfig(11 output_dir="dpo-out",12 beta=0.1,13 learning_rate=5e-7, # 10-100x lower than SFT. This matters.14 num_train_epochs=1,15 per_device_train_batch_size=2,16 gradient_accumulation_steps=8,17 max_length=1024, # prompt + response, in tokens18 loss_type="sigmoid", # also "ipo", "robust", "hinge", ...19 bf16=True,20 logging_steps=10,21)2223trainer = DPOTrainer(24 model=policy,25 ref_model=None, # with LoRA: adapters off == reference26 args=cfg,27 processing_class=tok,28 train_dataset=ds, # columns: prompt, chosen, rejected29 peft_config=LoraConfig(r=16, lora_alpha=32,30 target_modules=["q_proj", "v_proj"]),31)32trainer.train()Two details carry most of the practical weight. Setting ref_model=None alongside a LoRA config is not a shortcut — with adapters, the reference model is the base weights with adapters disabled, so it costs zero additional memory. And the learning rate is genuinely much lower than for SFT: DPO's gradient contains a difference of two log-likelihood gradients, which is a sharper signal than either alone, and 1e-5 will collapse the model within a few hundred steps.
DPO is a classification loss wearing reinforcement learning's clothes. Once the reward model cancels out, what remains is maximum likelihood on the winner minus maximum likelihood on the loser, weighted by how wrong you currently are.
Evaluating a DPO run
| Metric | Healthy pattern | Alarm |
|---|---|---|
rewards/accuracies | Rises from 0.50 to 0.65–0.80 | Stuck at 0.50 (bug) or above 0.95 (memorising) |
rewards/margins | Grows steadily, plateaus around 1–3 | Growing without bound — over-optimisation |
rewards/chosen | Slightly positive or near zero | Strongly negative alongside a positive margin |
rewards/rejected | Clearly negative | — |
| Mean generated length | Roughly stable | Climbing — length exploitation |
| Held-out win rate vs. SFT model | 60–75% | Below 55% — no real improvement |
The third row names a genuine and counter-intuitive property. DPO frequently drives the absolute log-probability of both chosen and rejected responses down, while widening the gap between them. The loss only cares about the difference, so pushing the rejected response down hard is a perfectly good way to satisfy it — and the chosen response can come along for the ride. Taken far enough, the model becomes less likely to produce the preferred answers themselves, and generations drift toward whatever the optimiser found in the gap. Watch rewards/chosen specifically: if it goes sharply negative, raise β or stop earlier.
When to use which
| Consideration | Favours DPO | Favours PPO |
|---|---|---|
| Preference data | Fixed dataset already collected | Ongoing collection; you can label the current policy's outputs |
| Compute | Limited — two models, no generation loop | Ample; generation infrastructure already in place |
| Reward signal | Only human comparisons | A programmatic reward exists (tests pass, verifier accepts) — DPO cannot use one; online RL such as PPO or GRPO can |
| Exploration | Not needed; the data defines the target | Needed; the model should discover responses better than anything in the dataset |
| Iteration speed | Hours to a first result | Days |
| Ceiling with abundant data and compute | Slightly lower in careful comparisons (Xu et al., 2024; Ivison et al., 2024) | Slightly higher — on-policy sampling keeps data matched to the model |
| Operational risk | Low; few things to misconfigure | High; many interacting hyperparameters |
The deepest difference is off-policy versus on-policy. DPO learns from a fixed dataset generated by some other model. As training proceeds, the policy moves away from whatever produced those responses, and the data becomes progressively less representative of what the model now does. PPO regenerates its data every step and never has this problem. The practical mitigation is iterative DPO: train, sample from the new model, collect fresh preferences on those samples, train again. Two or three rounds recover much of the gap while keeping the simple loss.
In practice the two are now often used in sequence rather than as rivals. Tülu 3 (Lambert et al., 2024), a fully open post-training recipe, runs SFT, then DPO on preference data, then RL with verifiable rewards on maths and instruction-following tasks where a checker can score the answer. DPO handles the fuzzy preferences cheaply; online RL handles what can be checked.
The variant family, briefly but usefully
| Method | Change | Problem it addresses |
|---|---|---|
| IPO (Azar et al., 2023) | Replaces the log-sigmoid with a squared loss targeting a fixed margin | DPO's loss keeps rewarding ever-larger margins, which drives over-optimisation on deterministic preferences |
| cDPO | Smooths labels toward ϵ instead of 0/1 | Noisy preference labels; stops the model pushing infinitely hard on a mislabelled pair |
| KTO (Ethayarajh et al., 2024) | Needs only "this output was good" or "this output was bad", not pairs | Real product feedback is thumbs-up/thumbs-down, not comparisons |
| ORPO (Hong et al., 2024) | Combines SFT and preference terms with an odds-ratio penalty; no reference model | Removes the separate SFT stage and the reference model entirely — one training run |
| SimPO (Meng et al., 2024) | Length-normalised implicit reward plus a target margin, no reference model | Length bias, plus the memory cost of the reference |
The pattern across all of them: DPO's loss makes three assumptions — that Bradley-Terry describes preferences, that labels are clean, and that summed sequence log-probability is the right quantity to compare — and each variant relaxes one of them.
In TRL 1.x, IPO and several other losses are simply loss_type values of DPOTrainer, KTO has its own KTOTrainer, and ORPO and SimPO (the latter via CPOTrainer with loss_type="simpo") live in trl.experimental. No variant wins everywhere; plain DPO with a well-chosen β remains a strong baseline, so try it first and switch only for the specific problem a variant addresses.
PPO regenerates its training data every step and never goes stale. DPO consumes a fixed dataset that its own progress makes less relevant — which is why the single most valuable thing you can add to a DPO pipeline is a second round.
What this means when you choose an approach
Start with DPO unless you have a specific reason not to. Two models, two hyperparameters, and a first result in an afternoon. If it gets you to an acceptable win rate, you have saved yourself an enormous amount of operational complexity. Reach for online RL (PPO, or a critic-free method such as GRPO) when you need on-policy exploration, when you have a programmatic reward that preference pairs cannot express, or when you have already exhausted what your fixed dataset can teach.
Check that the loss starts at 0.693. Every time. It is a one-line assertion that catches a whole family of silent bugs — unfrozen reference, mismatched checkpoints, wrong token masking — that otherwise present as a run that trains smoothly and produces nothing.
Report chosen and rejected lengths in every dataset you build. If chosen responses average 240 tokens and rejected ones average 130, your DPO run is going to learn "longer" as much as "better", and no hyperparameter fixes it downstream.
Track rewards/chosen, not just the margin. A growing margin with collapsing chosen reward is the signature failure of this method: the model is winning by suppressing the rejected response rather than by preferring the chosen one, and its generations degrade even as every logged number improves.
Plan for more than one round. A single pass over a fixed preference set is the cheapest useful thing you can do, and it is also the version of the method that leaves the most on the table. Sampling from the trained model, collecting fresh comparisons on those samples, and running again is where most of the remaining quality lives.