Course Content
Fine-Tuning LLMs
6 sections · 52 lessons
How does DPO work, and how does it differ from PPO-style RLHF?
What you need to know
The loss in words
margin = beta * [ (log pi(chosen) - log pi_ref(chosen)) - (log pi(rejected) - log pi_ref(rejected)) ]loss = -log sigmoid(margin)pi is the model being trained and pi_ref is a frozen copy of the starting (SFT) model. The loss goes down when the model raises the chosen answer's probability more than the reference does, and lowers the rejected answer's. The term beta * (log pi - log pi_ref) is called the implicit reward, which is where the paper's title, "Your Language Model is Secretly a Reward Model", comes from.
At step 0 the model equals the reference, so the margin is 0 and the loss is -log sigmoid(0) = ln 2 ≈ 0.693. If your first logged loss is not about 0.693, something is wrong with the setup.
DPO vs PPO
| DPO | PPO-style RLHF | |
|---|---|---|
| Models in memory | Policy + reference (the reference can be the base with the LoRA adapter switched off) | Policy, reference, reward, value |
| Training | Supervised loss, like SFT | Generate, score, update; many settings |
| Data | Fixed, offline preference pairs | Fresh answers generated during training |
| Exploration | None — only the pairs you gave | Can find better answers than in the data |
| Stability | High | Needs careful tuning |
Code (TRL 1.13)
1from trl import DPOConfig, DPOTrainer2from peft import LoraConfig34# each row: {"prompt": [...messages], "chosen": [...], "rejected": [...]}5args = DPOConfig(output_dir="dpo", beta=0.1, learning_rate=5e-6,6 num_train_epochs=1, per_device_train_batch_size=4)7trainer = DPOTrainer(model="my-org/brand-writer-sft", args=args,8 train_dataset=pref_ds,9 peft_config=LoraConfig(r=16, target_modules="all-linear",10 task_type="CAUSAL_LM"))11trainer.train()With peft_config, TRL uses the base model with the adapter disabled as the reference, so you do not load a second copy. DPO learning rates are small — around 1e-6 to 5e-6 — much lower than SFT.
What to watch
TRL logs rewards/margins, rewards/accuracies and logps/chosen. A known failure: DPO only cares about the gap, so it can lower the probability of both answers, the rejected one faster. If logps/chosen keeps falling, outputs can get worse even as the loss improves. Variants exist for this and for noisy labels, selected with loss_type — for example "ipo", or "robust", or label_smoothing for conservative DPO.
A real-life example
The fashion brand's writer is already fine-tuned (SFT) on approved descriptions, but some drafts are still bland. The brand team compares two drafts for 3,000 products:
- Chosen: "Hand-block printed in Sanganer on soft mulmul, this kurta stays cool through a Delhi June."
- Rejected: "This kurta is made of good quality fabric and is very comfortable to wear."
They run DPO on top of the SFT adapter with beta 0.1 for one epoch. Approval on held-out products rises, and they check logps/chosen to make sure the model has not simply become less likely to write anything fluent.
Follow-up questions to expect
- "What does beta do?" — It sets how strongly the model is held near the reference. Low beta allows big moves and more overfitting; high beta keeps it close. 0.1 is the usual start.
- "Where should chosen and rejected answers come from?" — Ideally from your own SFT model (on-policy). Pairs written by a different model teach the gap between two styles, not between good and bad answers from your model.
- "When would you still use online RL?" — When a checker can score fresh answers (maths, code, structured tasks), because online RL can explore beyond the dataset. GRPO is the common choice.