Reinforcement Learning from Human Feedback (RLHF)

Direct Preference Optimization (DPO)


Count the machinery a full PPO alignment run needs. A policy being trained. A frozen reference model. A reward model. A value model. Four sets of weights resident at once, generation inside the training loop, and a KL coefficient that will silently ruin the run if it is off by a factor of three. A single 7B pipeline can want 250 GB of accelerator memory before you have generated a token.

Now look at what that machinery ultimately produces. The reward model exists to turn preference pairs into a scalar. The scalar exists to guide the policy. The KL term exists to stop the policy leaving the region where the scalar is trustworthy. At the end you have a policy — and you throw the reward model away.

Direct Preference Optimization begins with a question: if the reward model is a disposable intermediate, is there a way to skip building it? The answer turns out to be yes, and not by approximation. Under the same assumptions PPO already makes, the language model is the reward model, expressed differently — and once you see that, the preference pairs can train the policy directly with an ordinary supervised loss.

PPO pipelineDPO pipeline
Models in memoryPolicy, reference, reward, valuePolicy, reference
Generation during trainingYes — dominates wall-clock timeNo
Training loopRollout, score, GAE, clipped updateOne forward pass, one loss, backprop
Key hyperparametersβ\beta, learning rate, clip range, GAE λ\lambda, epochs per rollout, value coefficientβ\beta, learning rate
Typical time to a first working runDaysHours
How the reward model cancels out1. RLHF maximises reward minus a KL penalty2. That objective has a closed-form optimal policy3. Invert it: reward is a log-ratio plus log Z4. Substituting into Bradley-Terry cancels log Z5. What remains is a loss on the policy alone
Nothing is approximated along the way — the reward model was always implicit in the policy, and step four is what makes it visible.

The derivation, one step at a time

Five steps. Each is short, and the payoff at step 4 is worth the effort of following steps 1 to 3.

Step 1 — write down what RLHF is optimising

The objective every KL-constrained RLHF method maximises is:

max⁡π  Ex∼D, y∼π(⋅∣x)[r(x,y)]  −  β DKL[π(y∣x) ∥ πref(y∣x)]\max_{\pi} \; \mathbb{E}_{x \sim \mathcal{D},\, y \sim \pi(\cdot \mid x)}\big[r(x,y)\big] \;-\; \beta\, \mathbb{D}_{\text{KL}}\big[\pi(y \mid x) \,\|\, \pi_{\text{ref}}(y \mid x)\big]

In words: get as much reward as possible, but pay a penalty proportional to how far the output distribution has moved from the frozen starting model. The coefficient β\beta prices that movement. PPO approximates a solution to this by sampling and taking constrained gradient steps.

Step 2 — this objective has an exact solution

You do not have to approximate it. Treating it as a constrained optimisation over distributions and solving gives a closed form:

π∗(y∣x)  =  1Z(x) πref(y∣x) exp⁡ ⁣(1βr(x,y)),Z(x)=∑yπref(y∣x) exp⁡ ⁣(1βr(x,y))\pi^*(y \mid x) \;=\; \frac{1}{Z(x)}\, \pi_{\text{ref}}(y \mid x)\, \exp\!\Big(\frac{1}{\beta} r(x,y)\Big), \qquad Z(x) = \sum_{y} \pi_{\text{ref}}(y \mid x)\, \exp\!\Big(\frac{1}{\beta} r(x,y)\Big)

Reading it: the optimal policy is the reference model, reweighted by the exponential of reward. Responses the reference already found plausible and that score well get boosted; responses that score badly get suppressed; β\beta controls how sharply. With β\beta large the exponent is small and π∗\pi^* barely differs from πref\pi_{\text{ref}}; with β\beta near zero it collapses onto whatever maximises reward.

This formula is exact but useless as an algorithm, because Z(x)Z(x) sums over every possible response — an astronomically large set. That intractability is precisely why PPO exists.

Step 3 — turn it inside out

Here is the move. That equation relates a policy to a reward. Solve it for the reward instead. Take logs of both sides and rearrange:

r(x,y)  =  βlog⁡π∗(y∣x)πref(y∣x)  +  βlog⁡Z(x)r(x,y) \;=\; \beta \log \frac{\pi^*(y \mid x)}{\pi_{\text{ref}}(y \mid x)} \;+\; \beta \log Z(x)

This says something surprising: any policy implicitly defines a reward function. The reward a policy "believes in" is β\beta times the log-ratio of its own probability to the reference's, plus a term that depends only on the prompt.

Z(x)Z(x) is still intractable. But notice it depends only on xx — the same value for every response to that prompt.

Step 4 — the cancellation

Preferences are modelled with Bradley-Terry, which says the probability a human prefers ywy_w to yly_l is a sigmoid of the reward difference:

P(yw≻yl∣x)=σ(r(x,yw)−r(x,yl))P(y_w \succ y_l \mid x) = \sigma\big(r(x, y_w) - r(x, y_l)\big)

Substitute the expression from step 3 for both rewards:

r(x,yw)−r(x,yl)=βlog⁡π∗(yw∣x)πref(yw∣x)+βlog⁡Z(x)−βlog⁡π∗(yl∣x)πref(yl∣x)−βlog⁡Z(x)r(x,y_w) - r(x,y_l) = \beta \log \frac{\pi^*(y_w \mid x)}{\pi_{\text{ref}}(y_w \mid x)} + \beta \log Z(x) - \beta \log \frac{\pi^*(y_l \mid x)}{\pi_{\text{ref}}(y_l \mid x)} - \beta \log Z(x)

The βlog⁡Z(x)\beta \log Z(x) terms are identical and cancel exactly. Because Bradley-Terry only ever sees differences, the intractable partition function vanishes — not approximated, not bounded, simply gone. The preference probability becomes:

P(yw≻yl∣x)=σ(βlog⁡π∗(yw∣x)πref(yw∣x)−βlog⁡π∗(yl∣x)πref(yl∣x))P(y_w \succ y_l \mid x) = \sigma\left(\beta \log \frac{\pi^*(y_w \mid x)}{\pi_{\text{ref}}(y_w \mid x)} - \beta \log \frac{\pi^*(y_l \mid x)}{\pi_{\text{ref}}(y_l \mid x)}\right)

The reward model was never fundamentally necessary. It was a way of parameterising something the policy and the reference already encode between them, and the only quantity that made it intractable cancels the moment you take a difference.

Step 5 — read it as a loss

Replace π∗\pi^* with the policy you are training, πθ\pi_\theta, and fit by maximum likelihood on your preference dataset:

LDPO(πθ;πref)=− E(x,yw,yl)∼D[log⁡σ ⁣(βlog⁡πθ(yw∣x)πref(yw∣x)−βlog⁡πθ(yl∣x)πref(yl∣x))]\mathcal{L}_{\text{DPO}}(\pi_\theta; \pi_{\text{ref}}) = -\,\mathbb{E}_{(x, y_w, y_l) \sim \mathcal{D}} \left[\log \sigma\!\left(\beta \log \frac{\pi_\theta(y_w \mid x)}{\pi_{\text{ref}}(y_w \mid x)} - \beta \log \frac{\pi_\theta(y_l \mid x)}{\pi_{\text{ref}}(y_l \mid x)}\right)\right]

Everything in it is computable with two forward passes. There is no sampling, no reward model, no value function, no rollouts. It is a classification loss over pairs.

Reading the loss without the notation

Define the implicit reward of a response as r^(x,y)=βlog⁡πθ(y∣x)πref(y∣x)\hat{r}(x,y) = \beta \log \frac{\pi_\theta(y \mid x)}{\pi_{\text{ref}}(y \mid x)} — how much more likely the policy makes this response than the reference did, scaled by β\beta. The loss is then simply −log⁡σ(r^w−r^l)-\log\sigma(\hat{r}_w - \hat{r}_l): make the policy relatively more enthusiastic about the winner than about the loser, compared with where it started.

The gradient makes the mechanism concrete:

∇θLDPO=−β E[σ(r^l−r^w)⏟how wrong we are(∇θlog⁡πθ(yw∣x)⏟push up−∇θlog⁡πθ(yl∣x)⏟push down)]\nabla_\theta \mathcal{L}_{\text{DPO}} = -\beta\, \mathbb{E}\Big[\underbrace{\sigma\big(\hat{r}_l - \hat{r}_w\big)}_{\text{how wrong we are}} \Big(\underbrace{\nabla_\theta \log \pi_\theta(y_w \mid x)}_{\text{push up}} - \underbrace{\nabla_\theta \log \pi_\theta(y_l \mid x)}_{\text{push down}}\Big)\Big]

Three things to notice. The bracketed part is ordinary maximum-likelihood on the winner minus maximum-likelihood on the loser — the same gradients supervised fine-tuning uses, with one sign flipped. The scalar in front is an automatic difficulty weighting: it is large when the model currently gets the pair backwards and small when the pair is already comfortably ranked. And the reference model appears only inside that weight, which is what keeps the policy tethered without any separate KL term.

A fully worked example

Take β=0.1\beta = 0.1 and one training pair. Sequence log-probabilities are sums over response tokens, so they are large negative numbers.

QuantityChosen ywy_wRejected yly_l
log⁡πθ(y∣x)\log \pi_\theta(y \mid x)-32.0-40.0
log⁡πref(y∣x)\log \pi_{\text{ref}}(y \mid x)-35.0-38.0
Log-ratio−32.0−(−35.0)=+3.0-32.0 - (-35.0) = +3.0−40.0−(−38.0)=−2.0-40.0 - (-38.0) = -2.0
Implicit reward r^=β×\hat{r} = \beta \times log-ratio0.1×3.0=+0.300.1 \times 3.0 = +0.300.1×(−2.0)=−0.200.1 \times (-2.0) = -0.20

Margin: r^w−r^l=0.30−(−0.20)=0.50\hat{r}_w - \hat{r}_l = 0.30 - (-0.20) = 0.50. Loss: −log⁡σ(0.50)=−log⁡(0.6225)=0.474-\log\sigma(0.50) = -\log(0.6225) = 0.474. Gradient weight: σ(−0.50)=0.3775\sigma(-0.50) = 0.3775.

Now the same arithmetic in three other regimes:

SituationMarginLossGradient weight σ(−margin)\sigma(-\text{margin})Interpretation
Step 0 — policy is an exact copy of the reference0.000.6930.500Both log-ratios are exactly zero. Every DPO run starts at loss log⁡2=0.693\log 2 = 0.693
Learning in progress0.500.4740.378Correct but not confident; substantial gradient
Pair mastered3.000.0490.047Gradient is 8x smaller than the previous row — the loss has stopped caring
Pair ranked backwards-1.501.7010.818Largest gradient in the batch; the update flips the ordering

That first row is the single most useful diagnostic in DPO. Your loss must start at 0.6931. If it does not, the policy and reference are not identical at initialisation — you loaded different checkpoints, or the reference is not frozen, or the log-probabilities are being computed over different token spans.

Why β\beta changes everything

Recompute the first example with different β\beta, holding the log-ratios at +3.0+3.0 and −2.0-2.0:

β\betar^w\hat{r}_wr^l\hat{r}_lMarginLossBehaviour
0.010.03-0.020.050.668Loss barely responds to a large probability shift — the policy can drift enormously before the loss objects. Degenerate outputs
0.10.30-0.200.500.474The standard default
0.51.50-1.002.500.079Loss is already nearly satisfied; the policy stays very close to the reference and learns little

β\beta in DPO plays exactly the role the KL coefficient plays in PPO: it is the price of moving away from the reference. Small β\beta means movement is cheap and the model drifts; large β\beta means movement is expensive and the model does not change. Sensible values sit between 0.05 and 0.5, with 0.1 the usual starting point.

Implementing it correctly

Python
import torchimport torch.nn.functional as Fdef sequence_logprob(model, input_ids, attention_mask, prompt_lens):    """Sum of log-probs over RESPONSE tokens only."""    logits = model(input_ids=input_ids,                   attention_mask=attention_mask).logits    # Shift: logits at position t predict the token at position t+1.    logits = logits[:, :-1, :]    labels = input_ids[:, 1:].clone()    # Mask out prompt tokens AND padding. Both are essential.    mask = attention_mask[:, 1:].clone().bool()    for i, plen in enumerate(prompt_lens):        mask[i, :max(plen - 1, 0)] = False    per_tok = torch.gather(        logits.log_softmax(-1), dim=2, index=labels.unsqueeze(2)    ).squeeze(2)    return (per_tok * mask).sum(-1)          # SUM, not meandef dpo_loss(policy, ref, batch, beta=0.1):    pol_w = sequence_logprob(policy, batch["w_ids"], batch["w_mask"],                             batch["prompt_lens"])    pol_l = sequence_logprob(policy, batch["l_ids"], batch["l_mask"],                             batch["prompt_lens"])    with torch.no_grad():                    # reference never trains        ref_w = sequence_logprob(ref, batch["w_ids"], batch["w_mask"],                                 batch["prompt_lens"])        ref_l = sequence_logprob(ref, batch["l_ids"], batch["l_mask"],                                 batch["prompt_lens"])    r_w = beta * (pol_w - ref_w)             # implicit rewards    r_l = beta * (pol_l - ref_l)    loss = -F.logsigmoid(r_w - r_l).mean()    return loss, {        "reward_accuracy": (r_w > r_l).float().mean().item(),        "reward_margin":   (r_w - r_l).mean().item(),        "reward_chosen":   r_w.mean().item(),        "reward_rejected": r_l.mean().item(),    }

Four lines in that function are where implementations go wrong.

MistakeWhat happensHow you spot it
Not masking prompt tokensThe shared prompt's log-probs are added to both sides. They mostly cancel in the difference, but the reference-model terms do not cancel cleanly and gradients flow into predicting the promptChosen and rejected implicit rewards both drift strongly in the same direction
Averaging per-token instead of summingYou have changed the objective. It is no longer the DPO loss, and length behaviour changes completely — this is a different algorithm (a reasonable one, but not this one)Loss does not start at 0.693 under a longer/shorter pair; results differ from published baselines
Reference model not in torch.no_grad() or not in eval()Gradients flow into the reference, or dropout makes reference log-probs stochastic, adding noise to every marginLoss is noisy at step 0 instead of exactly 0.693
Off-by-one in the logits/labels shiftEvery log-probability is for the wrong tokenReward accuracy hovers at 0.5 and never moves

The summing decision deserves elaboration because it drives DPO's best-known weakness. A 300-token response has a sequence log-probability roughly three times more negative than a 100-token one. Small per-token probability changes therefore produce much larger log-ratio swings on long responses. If your chosen responses are systematically longer than your rejected ones, DPO will find that increasing length is a cheap way to raise the margin — and you get the same length inflation that plagues reward models, arriving through a different door. Length-match your pairs where you can, and always report mean chosen and rejected lengths in your dataset statistics.

Preparing the data

DPO wants three columns — prompt, chosen, rejected — but the details of how you build them decide whether the run works.

Python
def build_pair(example, tok):    # 1. Format the prompt with the SAME chat template the SFT stage    #    used. A mismatch here is the most common cause of a run that    #    trains cleanly and produces a worse model: the policy is    #    being scored on a prompt format it was never trained on.    prompt = tok.apply_chat_template(        [{"role": "user", "content": example["question"]}],        tokenize=False, add_generation_prompt=True)    chosen   = example["chosen"]   + tok.eos_token    rejected = example["rejected"] + tok.eos_token    # 2. Tokenise prompt and response SEPARATELY, then concatenate,    #    so you know exactly where the prompt ends. Tokenising the    #    joined string can merge the boundary tokens and shift the    #    mask by one, which silently corrupts every log-probability.    p_ids = tok(prompt, add_special_tokens=False)["input_ids"]    c_ids = tok(chosen, add_special_tokens=False)["input_ids"]    r_ids = tok(rejected, add_special_tokens=False)["input_ids"]    # 3. Truncate the PROMPT from the left, never the response.    #    Cutting a response changes what you are comparing.    if len(p_ids) > MAX_PROMPT:        p_ids = p_ids[-MAX_PROMPT:]    return {"w_ids": p_ids + c_ids, "l_ids": p_ids + r_ids,            "prompt_len": len(p_ids)}

Three sanity checks before training, each of which takes a minute and each of which has saved runs: decode one w_ids and one l_ids back to text and read them, confirming the prompt appears exactly once and the response follows cleanly; assert that prompt_len is identical for the chosen and rejected members of every pair; and report the mean chosen and rejected response lengths across the dataset.

One more decision matters. DPO assumes the policy already assigns reasonable probability to responses like the chosen ones. If you run it directly on a base model that has never been instruction-tuned, both log-probabilities are tiny, the margins are dominated by noise, and training is unstable. The standard practice is to fine-tune on the chosen responses first — an ordinary SFT pass — and use that checkpoint as both the starting policy and the frozen reference.

Using TRL

Python
import torchfrom trl import DPOTrainer, DPOConfigfrom transformers import AutoModelForCausalLM, AutoTokenizerfrom peft import LoraConfigtok = AutoTokenizer.from_pretrained("sft-model")policy = AutoModelForCausalLM.from_pretrained("sft-model",                                              dtype=torch.bfloat16)cfg = DPOConfig(    output_dir="dpo-out",    beta=0.1,    learning_rate=5e-7,        # 10-100x lower than SFT. This matters.    num_train_epochs=1,    per_device_train_batch_size=2,    gradient_accumulation_steps=8,    max_length=1024,           # prompt + response, in tokens    loss_type="sigmoid",       # also "ipo", "robust", "hinge", ...    bf16=True,    logging_steps=10,)trainer = DPOTrainer(    model=policy,    ref_model=None,            # with LoRA: adapters off == reference    args=cfg,    processing_class=tok,    train_dataset=ds,          # columns: prompt, chosen, rejected    peft_config=LoraConfig(r=16, lora_alpha=32,                           target_modules=["q_proj", "v_proj"]),)trainer.train()

Two details carry most of the practical weight. Setting ref_model=None alongside a LoRA config is not a shortcut — with adapters, the reference model is the base weights with adapters disabled, so it costs zero additional memory. And the learning rate is genuinely much lower than for SFT: DPO's gradient contains a difference of two log-likelihood gradients, which is a sharper signal than either alone, and 1e-5 will collapse the model within a few hundred steps.

DPO is a classification loss wearing reinforcement learning's clothes. Once the reward model cancels out, what remains is maximum likelihood on the winner minus maximum likelihood on the loser, weighted by how wrong you currently are.

Evaluating a DPO run

MetricHealthy patternAlarm
rewards/accuraciesRises from 0.50 to 0.65–0.80Stuck at 0.50 (bug) or above 0.95 (memorising)
rewards/marginsGrows steadily, plateaus around 1–3Growing without bound — over-optimisation
rewards/chosenSlightly positive or near zeroStrongly negative alongside a positive margin
rewards/rejectedClearly negative—
Mean generated lengthRoughly stableClimbing — length exploitation
Held-out win rate vs. SFT model60–75%Below 55% — no real improvement

The third row names a genuine and counter-intuitive property. DPO frequently drives the absolute log-probability of both chosen and rejected responses down, while widening the gap between them. The loss only cares about the difference, so pushing the rejected response down hard is a perfectly good way to satisfy it — and the chosen response can come along for the ride. Taken far enough, the model becomes less likely to produce the preferred answers themselves, and generations drift toward whatever the optimiser found in the gap. Watch rewards/chosen specifically: if it goes sharply negative, raise β\beta or stop earlier.

When to use which

ConsiderationFavours DPOFavours PPO
Preference dataFixed dataset already collectedOngoing collection; you can label the current policy's outputs
ComputeLimited — two models, no generation loopAmple; generation infrastructure already in place
Reward signalOnly human comparisonsA programmatic reward exists (tests pass, verifier accepts) — DPO cannot use one; online RL such as PPO or GRPO can
ExplorationNot needed; the data defines the targetNeeded; the model should discover responses better than anything in the dataset
Iteration speedHours to a first resultDays
Ceiling with abundant data and computeSlightly lower in careful comparisons (Xu et al., 2024; Ivison et al., 2024)Slightly higher — on-policy sampling keeps data matched to the model
Operational riskLow; few things to misconfigureHigh; many interacting hyperparameters

The deepest difference is off-policy versus on-policy. DPO learns from a fixed dataset generated by some other model. As training proceeds, the policy moves away from whatever produced those responses, and the data becomes progressively less representative of what the model now does. PPO regenerates its data every step and never has this problem. The practical mitigation is iterative DPO: train, sample from the new model, collect fresh preferences on those samples, train again. Two or three rounds recover much of the gap while keeping the simple loss.

In practice the two are now often used in sequence rather than as rivals. Tülu 3 (Lambert et al., 2024), a fully open post-training recipe, runs SFT, then DPO on preference data, then RL with verifiable rewards on maths and instruction-following tasks where a checker can score the answer. DPO handles the fuzzy preferences cheaply; online RL handles what can be checked.

The variant family, briefly but usefully

MethodChangeProblem it addresses
IPO (Azar et al., 2023)Replaces the log-sigmoid with a squared loss targeting a fixed marginDPO's loss keeps rewarding ever-larger margins, which drives over-optimisation on deterministic preferences
cDPOSmooths labels toward ϵ\epsilon instead of 0/1Noisy preference labels; stops the model pushing infinitely hard on a mislabelled pair
KTO (Ethayarajh et al., 2024)Needs only "this output was good" or "this output was bad", not pairsReal product feedback is thumbs-up/thumbs-down, not comparisons
ORPO (Hong et al., 2024)Combines SFT and preference terms with an odds-ratio penalty; no reference modelRemoves the separate SFT stage and the reference model entirely — one training run
SimPO (Meng et al., 2024)Length-normalised implicit reward plus a target margin, no reference modelLength bias, plus the memory cost of the reference

The pattern across all of them: DPO's loss makes three assumptions — that Bradley-Terry describes preferences, that labels are clean, and that summed sequence log-probability is the right quantity to compare — and each variant relaxes one of them.

In TRL 1.x, IPO and several other losses are simply loss_type values of DPOTrainer, KTO has its own KTOTrainer, and ORPO and SimPO (the latter via CPOTrainer with loss_type="simpo") live in trl.experimental. No variant wins everywhere; plain DPO with a well-chosen β\beta remains a strong baseline, so try it first and switch only for the specific problem a variant addresses.

PPO regenerates its training data every step and never goes stale. DPO consumes a fixed dataset that its own progress makes less relevant — which is why the single most valuable thing you can add to a DPO pipeline is a second round.

What this means when you choose an approach

Start with DPO unless you have a specific reason not to. Two models, two hyperparameters, and a first result in an afternoon. If it gets you to an acceptable win rate, you have saved yourself an enormous amount of operational complexity. Reach for online RL (PPO, or a critic-free method such as GRPO) when you need on-policy exploration, when you have a programmatic reward that preference pairs cannot express, or when you have already exhausted what your fixed dataset can teach.

Check that the loss starts at 0.693. Every time. It is a one-line assertion that catches a whole family of silent bugs — unfrozen reference, mismatched checkpoints, wrong token masking — that otherwise present as a run that trains smoothly and produces nothing.

Report chosen and rejected lengths in every dataset you build. If chosen responses average 240 tokens and rejected ones average 130, your DPO run is going to learn "longer" as much as "better", and no hyperparameter fixes it downstream.

Track rewards/chosen, not just the margin. A growing margin with collapsing chosen reward is the signature failure of this method: the model is winning by suppressing the rejected response rather than by preferring the chosen one, and its generations degrade even as every logged number improves.

Plan for more than one round. A single pass over a fixed preference set is the cheapest useful thing you can do, and it is also the version of the method that leaves the most on the table. Sampling from the trained model, collecting fresh comparisons on those samples, and running again is where most of the remaining quality lives.