Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

What is ORPO and how does it streamline the usual SFT → preference-optimization workflow?


What you need to know

The loss

Text
L = L_SFT(chosen) + lambda * L_ORL_OR = -log sigmoid( log( odds(chosen) / odds(rejected) ) )

L_SFT teaches the model to write the chosen answer. L_OR adds the contrast: it is small only when the chosen answer's odds are clearly higher than the rejected answer's. The ORPO paper (Hong et al., 2024) observed that SFT alone also raises the probability of rejected-style answers, because they look similar; the odds term pushes them back down.

How it compares

MethodStagesReference modelData
SFT, then DPO2Yes, frozen (in DPO)SFT examples, then pairs
ORPO1NoPairs from the start
KTO1 (after SFT)YesSingle answers marked good or bad — no pairs
SimPO1 (after SFT)NoPairs; uses length-normalised rewards

Code (TRL 1.13)

Python
from trl.experimental.orpo import ORPOConfig, ORPOTrainerargs = ORPOConfig(output_dir="orpo", beta=0.1,   # beta is ORPO's lambda                  learning_rate=8e-6, num_train_epochs=1)trainer = ORPOTrainer(model="Qwen/Qwen2.5-7B", args=args,                      train_dataset=pairs_ds)   # prompt, chosen, rejectedtrainer.train()

In TRL 1.x, ORPO lives in trl.experimental, so its API may change between releases; DPO, KTO and GRPO are in the stable API.

A real-life example

The Hindi support team has a natural source of preference pairs: every time an agent edits the model's draft before sending it, the edited reply is "chosen" and the original draft is "rejected". After three months they have 5,000 such pairs.

They compare two pipelines starting from the base model: SFT on the edited replies then DPO, versus one ORPO run on the pairs. On 300 fresh chats rated by senior agents, the two score within a point of each other. ORPO took one run instead of two and needed no reference model in memory, so the team adopts it for monthly retraining — while keeping DPO as a fallback if a future ORPO release changes behaviour.

Follow-up questions to expect

  • "Why no reference model?" — The odds-ratio term compares chosen and rejected under the current model only, and the SFT term keeps the model anchored to good answers, so there is no need for a KL constraint against a frozen copy.
  • "When is KTO better?" — When you only have thumbs-up or thumbs-down on single answers, not pairs. It is common for feedback collected in a product.
  • "What does the lambda (beta) weight do?" — It balances imitation against contrast. Too high and the model focuses on beating rejected answers at the expense of fluent writing; too low and it is plain SFT.