Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

What is RLAIF (Reinforcement Learning from AI Feedback), and how is it different from RLHF?


What you need to know

RLHFRLAIF
Who labelsPeopleAn LLM judge
Cost and speedHigh; days to weeksLow; hours
ScaleThousands of pairsUp to millions
ConsistencyVaries by person and fatigueVery consistent — including consistently wrong
Where bias comes fromLabeller backgrounds, unclear guidesThe judge model's own preferences

Research from Google in 2023 (Lee et al.) reported RLAIF performing comparably to RLHF on tasks such as summarisation, which is why it became common. But the judge sets the ceiling.

Known judge biases

  • Position bias — preferring whichever answer is shown first (or second).
  • Length bias — preferring longer, more detailed-looking answers.
  • Self-preference — preferring text written in its own style or by its own model family.

Making the judge trustworthy

Python
def judge_pair(prompt, a, b):    first = judge(prompt, answer_1=a, answer_2=b)   # returns "1", "2" or "tie"    second = judge(prompt, answer_1=b, answer_2=a)  # same pair, order swapped    if first == "1" and second == "2":        return "a"    if first == "2" and second == "1":        return "b"    return "tie"   # inconsistent verdict: drop from preference data

Asking twice with the order swapped removes position bias: only consistent verdicts become training pairs. Also ask the judge to reason against the rubric before its verdict, use a judge from a different model family than the policy, and measure its agreement with human labels on a few hundred pairs.

A real-life example

The fashion brand wants DPO data for 50,000 product descriptions — far more than its three-person brand team can label. It writes a nine-rule brand constitution ("name the fabric and craft", "no unverifiable claims like 'eco-friendly' unless the catalogue says so", "no clichés") and uses a strong model as judge with swapped-order checks.

On 500 pairs the brand team also labelled, the judge agrees with them 82% of the time overall — but only 60% on descriptions that mention sustainability, where a wrong claim is a legal risk. The team keeps the AI labels for everything else and sends sustainability-related pairs to humans. Cheap volume where the judge is reliable, human labels where it is not.

Follow-up questions to expect

  • "Can the judge score be used directly as the RL reward?" — Yes; that skips training a separate reward model. It costs a judge call per sample, and the policy can learn to exploit the judge's quirks, so keep checking against humans.
  • "Is RLAIF the same as distillation?" — No. Distillation copies a teacher's answers; RLAIF uses a model only to compare answers, and the policy learns from those preferences.
  • "When must humans stay in the loop?" — Safety, legal and culturally sensitive categories, and whenever judge–human agreement is low for a category.