Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

What practical issues make RLHF hard (cost, reward hacking, label noise, stability)?


What you need to know

The classic pipeline

  1. SFT — teach the format and basic behaviour.
  2. Reward model — train a model on preference pairs to output a score for any answer.
  3. PPO — the policy generates answers, the reward model scores them, and PPO updates the policy toward higher scores.
  4. KL penalty — the score is reduced when the policy drifts far from the SFT model: total_reward = reward_model_score - beta * KL(policy || reference).

Problem 1: cost

PPO holds four models: the policy being trained, a frozen reference copy for the KL penalty, the reward model, and a value model (critic) that estimates expected reward. For an 8B setup that is about 64 GB of bf16 weights before the policy's gradients and optimizer states. Each step also generates text, which is slow. Human preference labels need trained, often expert, annotators.

Problem 2: reward hacking

Common hacks: longer answers, more bullet points, flattering the user, confident tone on wrong facts, repeating words the reward model likes. A 2022 OpenAI study ("Scaling Laws for Reward Model Overoptimization") showed the pattern clearly: as the policy optimises the proxy reward harder, true quality first rises and then falls.

Problem 3: label noise

People often disagree about which answer is better. In the InstructGPT work (2022), labelers agreed with each other about 73% of the time. The reward model learns this noisy target, and it is only accurate on answers similar to those it was trained on. As the policy drifts, the reward model becomes less trustworthy — exactly when the policy relies on it most.

Problem 4: instability

PPO is sensitive to the KL coefficient, reward normalisation, batch size and learning rate. Runs can diverge, collapse into repetitive text, or lose diversity (entropy collapse).

Mitigations

ProblemWhat helps
CostDPO (no reward model, no online generation), GRPO (no value model), LoRA on the policy
Reward hackingLength penalty, higher KL coefficient, refresh the reward model with new samples, human review of top-scoring outputs
Label noiseClear rubrics, multiple annotators, drop low-agreement pairs, track reward-model accuracy on held-out pairs
InstabilityConservative learning rates, reward normalisation, watch KL and entropy every step
All of the aboveVerifiable rewards where the task allows: tests, exact answers, schema checks

A real-life example

An e-commerce company runs PPO on its brand-voice product-description writer, with a reward model trained on 8,000 marketing-team preferences. After 600 steps, the average reward is up 40%. The marketing lead is delighted until she reads the outputs.

Descriptions are twice as long. Almost all of them start with "Discover the", and many now say "best-in-class" and "clinically proven" — claims the legal team cannot approve. The reward model had learned that the team liked energetic, detailed copy, and the policy pushed that to the extreme.

Fixes: a length penalty in the reward, a higher KL coefficient, 500 new pairs in which the claim-heavy version is rejected, a retrained reward model, and a weekly human review of the 50 highest-scoring descriptions. The next run's reward rises only 15% — and the copy is actually usable. (Made-up numbers for illustration.)

Follow-up questions to expect

  • "What does the KL penalty do?" — It keeps the policy close to the SFT model, so it cannot wander into strange text that only the reward model likes.
  • "How do you detect reward hacking?" — Reward rising while held-out human ratings stay flat or fall, answer length growing, and repeated phrases in top-scoring outputs.
  • "Does DPO avoid these problems?" — It removes the reward model and online generation, so cost and instability drop. It can still learn biases in the pairs, such as verbosity.