Course Content
Fine-Tuning LLMs
6 sections · 52 lessons
Beyond plain SFT, what other fine-tuning approaches are commonly used?
What you need to know
SFT (supervised fine-tuning) shows the model one good answer per prompt and trains it to copy that answer. It cannot say "this other answer is worse", and it can only be as good as the examples you have. Every other method adds something SFT lacks.
The main options, grouped by signal
| Method | Signal it learns from | Data you need | Cost and complexity |
|---|---|---|---|
| SFT | One good answer | Prompt and answer pairs | Low |
| Rejection sampling (best-of-n) | Model's own answers that pass a checker | Prompts plus a checker or judge | Low: generate, filter, SFT |
| DPO, IPO, SimPO, ORPO | Better versus worse answer | Prompt, chosen, rejected | Low to medium; trains like SFT |
| KTO | Single thumbs up or down | Prompt, answer, good/bad flag | Low to medium |
| PPO (classic RLHF) | Score from a learned reward model | Preference pairs, then prompts | High: four models, online generation |
| GRPO with verifiable rewards | Score from code: tests, exact answer, schema | Prompts plus a reward function | Medium to high |
| Continual pretraining | Next token on raw domain text | Millions to billions of tokens | Medium to high |
| Distillation | A bigger teacher's outputs or probabilities | Prompts plus teacher access | Medium |
GRPO in one paragraph
With verifiable rewards (RLVR), the score comes from code instead of a learned model, so there is far less to hack:
1# trl 1.x2from trl import GRPOConfig, GRPOTrainer34def label_reward(completions, label, **kwargs):5 # completions are chat messages; `label` is a column from the dataset6 preds = [c[0]["content"].strip().split()[-1].lower() for c in completions]7 return [1.0 if p == g else 0.0 for p, g in zip(preds, label)]89args = GRPOConfig(output_dir="clause-grpo", num_generations=8, learning_rate=1e-6)10trainer = GRPOTrainer(model="./clause-sft-merged", reward_funcs=label_reward,11 args=args, train_dataset=prompts_with_labels)The reward function gets every generated answer plus the dataset columns, and returns one number per answer.
Where each lives today
In TRL 1.x, SFTTrainer, DPOTrainer, KTOTrainer, GRPOTrainer, RLOOTrainer and RewardTrainer are in the main package; ORPO and several newer methods sit in trl.experimental. Hosted APIs also offer more than SFT: OpenAI's fine-tuning API, for example, accepts supervised, dpo and reinforcement methods.
A real-life example
A hospital chain has an SFT-trained medical-report summariser. Doctors say it sometimes adds a number that is not in the report — the worst possible error.
- Rejection sampling. For 5,000 reports, they generate 4 summaries each and keep only those where every number and drug name also appears in the source report (a simple code check). They fine-tune on the survivors. Invented numbers drop sharply, with no new human labels.
- DPO. Doctors' edits to live summaries become pairs: the edited summary is
chosen, the original isrejected. 3,000 pairs later, the summaries match the doctors' preferred structure. - No PPO. There is no reliable reward model for "clinically useful", so full RLHF is not worth the cost.
Follow-up questions to expect
- "When would you choose GRPO over DPO?" — When you can compute a trustworthy score in code (tests pass, answer matches, JSON validates) and want the model to explore beyond your existing examples. DPO needs only pairs and is simpler.
- "What is KTO good for?" — Data that has only thumbs-up or thumbs-down on single answers, which is what most production feedback looks like.
- "Isn't rejection sampling just SFT?" — The training step is SFT, but the data is the model's own filtered outputs, which improves it without new human labels.