Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

How does DPO work, and how does it differ from PPO-style RLHF?


What you need to know

The loss in words

Text
margin = beta * [ (log pi(chosen)   - log pi_ref(chosen))                - (log pi(rejected) - log pi_ref(rejected)) ]loss   = -log sigmoid(margin)

pi is the model being trained and pi_ref is a frozen copy of the starting (SFT) model. The loss goes down when the model raises the chosen answer's probability more than the reference does, and lowers the rejected answer's. The term beta * (log pi - log pi_ref) is called the implicit reward, which is where the paper's title, "Your Language Model is Secretly a Reward Model", comes from.

At step 0 the model equals the reference, so the margin is 0 and the loss is -log sigmoid(0) = ln 2 ≈ 0.693. If your first logged loss is not about 0.693, something is wrong with the setup.

DPO vs PPO

DPOPPO-style RLHF
Models in memoryPolicy + reference (the reference can be the base with the LoRA adapter switched off)Policy, reference, reward, value
TrainingSupervised loss, like SFTGenerate, score, update; many settings
DataFixed, offline preference pairsFresh answers generated during training
ExplorationNone — only the pairs you gaveCan find better answers than in the data
StabilityHighNeeds careful tuning

Code (TRL 1.13)

Python
from trl import DPOConfig, DPOTrainerfrom peft import LoraConfig# each row: {"prompt": [...messages], "chosen": [...], "rejected": [...]}args = DPOConfig(output_dir="dpo", beta=0.1, learning_rate=5e-6,                 num_train_epochs=1, per_device_train_batch_size=4)trainer = DPOTrainer(model="my-org/brand-writer-sft", args=args,                     train_dataset=pref_ds,                     peft_config=LoraConfig(r=16, target_modules="all-linear",                                                 task_type="CAUSAL_LM"))trainer.train()

With peft_config, TRL uses the base model with the adapter disabled as the reference, so you do not load a second copy. DPO learning rates are small — around 1e-6 to 5e-6 — much lower than SFT.

What to watch

TRL logs rewards/margins, rewards/accuracies and logps/chosen. A known failure: DPO only cares about the gap, so it can lower the probability of both answers, the rejected one faster. If logps/chosen keeps falling, outputs can get worse even as the loss improves. Variants exist for this and for noisy labels, selected with loss_type — for example "ipo", or "robust", or label_smoothing for conservative DPO.

A real-life example

The fashion brand's writer is already fine-tuned (SFT) on approved descriptions, but some drafts are still bland. The brand team compares two drafts for 3,000 products:

  • Chosen: "Hand-block printed in Sanganer on soft mulmul, this kurta stays cool through a Delhi June."
  • Rejected: "This kurta is made of good quality fabric and is very comfortable to wear."

They run DPO on top of the SFT adapter with beta 0.1 for one epoch. Approval on held-out products rises, and they check logps/chosen to make sure the model has not simply become less likely to write anything fluent.

Follow-up questions to expect

  • "What does beta do?" — It sets how strongly the model is held near the reference. Low beta allows big moves and more overfitting; high beta keeps it close. 0.1 is the usual start.
  • "Where should chosen and rejected answers come from?" — Ideally from your own SFT model (on-policy). Pairs written by a different model teach the gap between two styles, not between good and bad answers from your model.
  • "When would you still use online RL?" — When a checker can score fresh answers (maths, code, structured tasks), because online RL can explore beyond the dataset. GRPO is the common choice.