Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

Beyond plain SFT, what other fine-tuning approaches are commonly used?


What you need to know

SFT (supervised fine-tuning) shows the model one good answer per prompt and trains it to copy that answer. It cannot say "this other answer is worse", and it can only be as good as the examples you have. Every other method adds something SFT lacks.

The main options, grouped by signal

MethodSignal it learns fromData you needCost and complexity
SFTOne good answerPrompt and answer pairsLow
Rejection sampling (best-of-n)Model's own answers that pass a checkerPrompts plus a checker or judgeLow: generate, filter, SFT
DPO, IPO, SimPO, ORPOBetter versus worse answerPrompt, chosen, rejectedLow to medium; trains like SFT
KTOSingle thumbs up or downPrompt, answer, good/bad flagLow to medium
PPO (classic RLHF)Score from a learned reward modelPreference pairs, then promptsHigh: four models, online generation
GRPO with verifiable rewardsScore from code: tests, exact answer, schemaPrompts plus a reward functionMedium to high
Continual pretrainingNext token on raw domain textMillions to billions of tokensMedium to high
DistillationA bigger teacher's outputs or probabilitiesPrompts plus teacher accessMedium

GRPO in one paragraph

With verifiable rewards (RLVR), the score comes from code instead of a learned model, so there is far less to hack:

Python
# trl 1.xfrom trl import GRPOConfig, GRPOTrainerdef label_reward(completions, label, **kwargs):    # completions are chat messages; `label` is a column from the dataset    preds = [c[0]["content"].strip().split()[-1].lower() for c in completions]    return [1.0 if p == g else 0.0 for p, g in zip(preds, label)]args = GRPOConfig(output_dir="clause-grpo", num_generations=8, learning_rate=1e-6)trainer = GRPOTrainer(model="./clause-sft-merged", reward_funcs=label_reward,                      args=args, train_dataset=prompts_with_labels)

The reward function gets every generated answer plus the dataset columns, and returns one number per answer.

Where each lives today

In TRL 1.x, SFTTrainer, DPOTrainer, KTOTrainer, GRPOTrainer, RLOOTrainer and RewardTrainer are in the main package; ORPO and several newer methods sit in trl.experimental. Hosted APIs also offer more than SFT: OpenAI's fine-tuning API, for example, accepts supervised, dpo and reinforcement methods.

A real-life example

A hospital chain has an SFT-trained medical-report summariser. Doctors say it sometimes adds a number that is not in the report — the worst possible error.

  1. Rejection sampling. For 5,000 reports, they generate 4 summaries each and keep only those where every number and drug name also appears in the source report (a simple code check). They fine-tune on the survivors. Invented numbers drop sharply, with no new human labels.
  2. DPO. Doctors' edits to live summaries become pairs: the edited summary is chosen, the original is rejected. 3,000 pairs later, the summaries match the doctors' preferred structure.
  3. No PPO. There is no reliable reward model for "clinically useful", so full RLHF is not worth the cost.

Follow-up questions to expect

  • "When would you choose GRPO over DPO?" — When you can compute a trustworthy score in code (tests pass, answer matches, JSON validates) and want the model to explore beyond your existing examples. DPO needs only pairs and is simpler.
  • "What is KTO good for?" — Data that has only thumbs-up or thumbs-down on single answers, which is what most production feedback looks like.
  • "Isn't rejection sampling just SFT?" — The training step is SFT, but the data is the model's own filtered outputs, which improves it without new human labels.