Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

If preference labels disagree (low annotator agreement), how do you improve RLHF/DPO data quality?


What you need to know

Worked example

Two raters label 200 pairs as "A better" or "B better". They agree on 140 (70%). If each picks A about half the time, chance agreement is 50%:

Text
kappa = (0.70 - 0.50) / (1 - 0.50) = 0.40

That is only "moderate" agreement. Below about 0.3, the task itself is probably not well defined.

Fix the rubric

Replace "which answer is better?" with ordered criteria, so raters stop optimising different things:

  1. Factually correct and consistent with policy.
  2. Actually solves the user's problem.
  3. Follows the instructions and matches the user's language.
  4. Brief and polite.

Also collect a rating per criterion, not only an overall choice. You learn where raters disagree, and you can combine the ratings later.

Fix the process

  • Gold set — 50–100 pairs with agreed answers; onboard raters on it and re-check them regularly.
  • More raters for hard items — 3–5 raters on ambiguous prompts; one rater on easy ones to save cost.
  • Allow ties — forcing a choice between equal answers creates coin-flip labels. Drop ties from DPO data.
  • Keep clear margins — pairs where raters strongly agree carry signal; near-ties mostly add noise.
  • Use an LLM judge as an extra rater — pairs where humans and the judge disagree often point to rubric gaps.

Robust training options in TRL 1.13

Python
from trl import DPOConfigargs = DPOConfig(output_dir="dpo", beta=0.1,                 label_smoothing=0.1)   # assume ~10% of labels are flipped# alternatives: loss_type=["ipo"] or loss_type=["robust"]

label_smoothing gives conservative DPO: the loss assumes some labels are wrong and does not push to extremes. IPO was designed to avoid overfitting to preferences that look certain. These help with leftover noise; they do not rescue a broken rubric.

A real-life example

The Hindi support team double-labels 400 reply pairs and gets a kappa of 0.28. Reading the disagreements shows two camps: some agents prefer formal Hindi with "aap" and full sentences; others prefer short casual Hinglish that matches the customer. Some agents also rate speed of resolution above tone.

The team writes the four ordered criteria above, with "match the customer's language and register" made explicit, runs a one-hour calibration on 60 gold pairs, and relabels. Kappa rises to 0.61. They drop ties (about 12% of pairs) and train DPO with label_smoothing=0.05. The new model's replies are rated better on a fresh human test set than the model trained on the original labels.

Follow-up questions to expect

  • "What kappa is good enough?" — Above about 0.6 is usually workable for preference data; below 0.4, fix the rubric before collecting more.
  • "Should you average ratings from several raters?" — Use majority vote or confidence-weighted votes, and consider dropping pairs with split votes.
  • "Can a reward model handle noisy labels?" — Somewhat, since it learns an average preference, but systematic disagreement (two camps) produces a reward model that satisfies neither.