Course Content
Fine-Tuning LLMs
6 sections · 52 lessons
What is ORPO and how does it streamline the usual SFT → preference-optimization workflow?
What you need to know
The loss
L = L_SFT(chosen) + lambda * L_ORL_OR = -log sigmoid( log( odds(chosen) / odds(rejected) ) )L_SFT teaches the model to write the chosen answer. L_OR adds the contrast: it is small only when the chosen answer's odds are clearly higher than the rejected answer's. The ORPO paper (Hong et al., 2024) observed that SFT alone also raises the probability of rejected-style answers, because they look similar; the odds term pushes them back down.
How it compares
| Method | Stages | Reference model | Data |
|---|---|---|---|
| SFT, then DPO | 2 | Yes, frozen (in DPO) | SFT examples, then pairs |
| ORPO | 1 | No | Pairs from the start |
| KTO | 1 (after SFT) | Yes | Single answers marked good or bad — no pairs |
| SimPO | 1 (after SFT) | No | Pairs; uses length-normalised rewards |
Code (TRL 1.13)
1from trl.experimental.orpo import ORPOConfig, ORPOTrainer23args = ORPOConfig(output_dir="orpo", beta=0.1, # beta is ORPO's lambda4 learning_rate=8e-6, num_train_epochs=1)5trainer = ORPOTrainer(model="Qwen/Qwen2.5-7B", args=args,6 train_dataset=pairs_ds) # prompt, chosen, rejected7trainer.train()In TRL 1.x, ORPO lives in trl.experimental, so its API may change between releases; DPO, KTO and GRPO are in the stable API.
A real-life example
The Hindi support team has a natural source of preference pairs: every time an agent edits the model's draft before sending it, the edited reply is "chosen" and the original draft is "rejected". After three months they have 5,000 such pairs.
They compare two pipelines starting from the base model: SFT on the edited replies then DPO, versus one ORPO run on the pairs. On 300 fresh chats rated by senior agents, the two score within a point of each other. ORPO took one run instead of two and needed no reference model in memory, so the team adopts it for monthly retraining — while keeping DPO as a fallback if a future ORPO release changes behaviour.
Follow-up questions to expect
- "Why no reference model?" — The odds-ratio term compares chosen and rejected under the current model only, and the SFT term keeps the model anchored to good answers, so there is no need for a KL constraint against a frozen copy.
- "When is KTO better?" — When you only have thumbs-up or thumbs-down on single answers, not pairs. It is common for feedback collected in a product.
- "What does the lambda (beta) weight do?" — It balances imitation against contrast. Too high and the model focuses on beating rejected answers at the expense of fluent writing; too low and it is plain SFT.