Course Content
Fine-Tuning LLMs
6 sections · 52 lessons
What is RLHF and how is it used to align an LLM’s behavior with human preferences?
What you need to know
The classic three stages
- SFT — fine-tune on demonstrations so the model follows instructions at all.
- Reward model — for each prompt, show people two answers and ask which is better. Train a model so the chosen answer gets the higher score.
- RL with PPO — the policy writes answers, the reward model scores them, and PPO updates the policy to raise
reward - beta * KL.
The reward model uses a pairwise loss:
loss = -log sigmoid( r(prompt, chosen) - r(prompt, rejected) )It only learns that "chosen" should score higher than "rejected", not an absolute quality number.
Why PPO is heavy
PPO keeps four models in play: the policy, a frozen reference for the KL penalty, the reward model, and a value model (critic) that predicts expected reward so updates are less noisy. It also needs text generation inside the training loop. That is a lot of memory and many settings to tune.
Reward hacking
The policy optimises the reward model, not true human preference. If the reward model slightly prefers longer or more flattering answers, the policy learns to be long and flattering. This is reward hacking, and the KL penalty only limits it.
What teams use in 2026
- DPO — learns from preference pairs directly, with no reward model or RL loop (next lesson).
- GRPO — introduced in DeepSeekMath (2024) and used for DeepSeek-R1 (2025). It drops the value model: for each prompt it samples a group of answers, for example 8, scores them, and uses each answer's score relative to the group average as its advantage.
- RL with verifiable rewards (RLVR) — for maths and code, a checker (exact answer match, unit tests) replaces the learned reward model, which removes most reward hacking on those tasks.
In TRL 1.13, GRPOTrainer and RLOOTrainer are the online RL trainers. TRL no longer ships a PPOTrainer; large PPO runs use frameworks such as OpenRLHF or veRL.
1from trl import GRPOConfig, GRPOTrainer23def correct_total(completions, answer, **kwargs):4 # 1.0 if the reply ends with the right total, else 0.05 return [1.0 if c[0]["content"].strip().endswith(a) else 0.06 for c, a in zip(completions, answer)]78args = GRPOConfig(output_dir="grpo", num_generations=8, learning_rate=1e-6)9trainer = GRPOTrainer(model="Qwen/Qwen3-1.7B", reward_funcs=correct_total,10 args=args, train_dataset=train_ds) # columns: prompt, answerExtra dataset columns such as answer are passed to the reward function by name. If all 8 answers in a group get the same score, their advantages are zero and that prompt teaches nothing — so prompts must be neither too easy nor too hard.
A real-life example
The Hindi support team collects 4,000 comparisons: senior agents see two replies to the same customer message and pick the better one. They train a reward model and run RL.
After a few hundred steps the reward keeps rising, but the weekly customer-satisfaction score does not. Reading samples shows why: replies have grown 60% longer and start with two sentences of apology. The agents had slightly preferred polite replies, and the policy found that more apology meant more reward. The team adds a length penalty, rewrites the rating guide to say "polite but brief", and relabels 500 pairs.
Follow-up questions to expect
- "Why is the KL penalty needed?" — Without it, the policy drifts toward odd text that the reward model happens to score highly. The penalty keeps it near the SFT model, where the reward model's scores are trustworthy.
- "Why compare two answers instead of scoring one?" — People are more consistent at "A is better than B" than at "this is a 7 out of 10".
- "How is GRPO different from PPO?" — GRPO has no value model; it compares each answer to others in its group. That saves memory and suits tasks with a checkable reward. Recent TRL versions even default GRPO's KL coefficient to 0.