Course Content
Reinforcement Learning from Human Feedback (RLHF)
4 sections · 10 lessons
What is RLHF and Why It's Needed
In the work that became InstructGPT (Ouyang et al., 2022), OpenAI researchers gave GPT-3 this prompt: "Explain the moon landing to a six year old in a few sentences." The model, trained on hundreds of billions of words of internet text, replied:
Explain the theory of gravity to a 6 year old. Explain the theory of relativity to a 6 year old in a few sentences. Explain the big bang theory to a 6 year old.
Nothing is wrong with the model. It did exactly what it was trained to do. Its training objective was predict the next token over a corpus scraped from the web, and on the web, a line that looks like a homework prompt is usually followed by more homework prompts, not by an answer. The model produced the statistically likely continuation. It was not being unhelpful; it had never been told that "helpful" was the goal.
This is the gap that Reinforcement Learning from Human Feedback exists to close. The gap is not one of capability — the same GPT-3 could explain the moon landing beautifully if you set the prompt up right. The gap is between what the training objective rewarded and what humans actually wanted. RLHF is a family of techniques for taking human judgements about which outputs are better, compressing them into a learned scoring function, and then using that function to reshape the model's behaviour.
The headline result from that same research: a 1.3-billion-parameter model tuned with RLHF was preferred by human raters over the original 175-billion-parameter GPT-3, a model over 100 times larger. Alignment bought more perceived quality than a 100x increase in scale.
First, try the obvious fix and watch it fail
Before reaching for reinforcement learning, try the straightforward thing. If the model does not know that questions should be answered, show it examples of questions being answered. Hire writers, have them compose thousands of ideal responses, and fine-tune on those pairs with the usual cross-entropy loss. This is supervised fine-tuning (SFT), sometimes called instruction tuning, and it genuinely works. It is the first stage of essentially every aligned model shipped today.
But run it to its limit and three walls appear.
Wall 1: writing is far more expensive than judging
A good demonstration of "summarise this 900-word article well" takes a skilled annotator around 15 minutes: read, draft, revise. Judging which of two candidate summaries is better takes about 30 seconds. That is a 30x difference in throughput from the same person.
Put real numbers on it. Suppose annotators cost 25 units of currency per hour.
| Approach | Time per item | Items per annotator-hour | Cost per 50,000 items |
|---|---|---|---|
| Write an ideal response | 15 min | 4 | 312,500 |
| Compare two responses | 30 sec | 120 | 10,417 |
Thirty times cheaper for the same annotator budget. If your budget is fixed, comparison data gives you thirty times more signal.
Wall 2: cross-entropy scores tokens, not answers
Supervised fine-tuning maximises the log-probability of the reference answer, token by token:
Reading it plainly: for each position t in the target answer, ask the model how much probability it assigned to the token that the human actually wrote, take the logarithm, and push it up. The problem is what this treats as an error. A response that is factually perfect but phrased differently from the reference gets penalised on almost every token. A response that copies the reference's phrasing but inverts one crucial fact gets penalised on roughly one token. The loss has no notion of "this answer is good" — only "this answer is worded like the reference".
Supervised fine-tuning teaches a model to imitate a particular answer. It cannot teach a model that one answer is better than another, because it never sees a worse answer to contrast against.
Wall 3: you cannot demonstrate what you cannot write
Ask an annotator to write the ideal response to "Should I take out a loan to invest in cryptocurrency?" They need to be helpful without giving financial advice, acknowledge risk without being preachy, and stay honest about uncertainty. Ten thoughtful annotators produce ten different answers, and each would find flaws in the others'. There is no single target to imitate.
Yet show those same ten annotators two candidate responses — one that says "absolutely, crypto is going up" and one that lays out the risk of leveraged speculation — and nine or ten will pick the same one immediately. Judgement is reliable where authorship is not.
This asymmetry is the entire foundation of RLHF. Humans are far better at recognising quality than at producing it on demand.
The shape of the pipeline
RLHF splits the problem into three stages, each solving a piece the others cannot.
Stage 0 Pretrained base model "knows language, doesn't know it should be helpful" | vStage 1 Supervised fine-tuning (SFT) ~10k-100k human-written demonstrations teaches the FORMAT: answer questions, follow instructions | vStage 2 Reward model (RM) training ~50k-1M human preference comparisons (A vs B) teaches a SCORER: r(prompt, response) -> scalar | vStage 3 Policy optimisation (PPO, DPO, ...) model generates -> RM scores -> model updated to score higher constrained so it doesn't drift far from Stage 1 | v Aligned modelNotice the division of labour. Stage 1 fixes the format problem — after SFT, the model at least tries to answer. Stage 2 turns scattered human opinions into a differentiable function you can evaluate millions of times without asking a human. Stage 3 uses that function as a training signal.
Why not skip stage 2 and let humans score the model's outputs directly during training? Because policy optimisation needs on the order of hundreds of thousands to millions of scored samples, generated on the fly, changing every few minutes as the policy shifts. No annotation pipeline runs at that rate. The reward model is a cheap, fast, differentiable stand-in for a human rater — that is its whole job, and, as we will see, also its whole weakness.
The reward model and the Bradley-Terry assumption
Your raw data is comparisons: for prompt x, response yw was preferred over response yl ("w" for winner, "l" for loser). You want a function rϕ(x,y) giving a single number — higher is better. How do you turn "A beat B" into a numeric score?
The standard answer borrows from a 1952 model of paired comparisons in sports rankings, the Bradley-Terry model. It assumes every option has a latent strength, and the probability that one beats another depends only on the difference in strengths:
Read this in plain English: the chance a human prefers response A over response B is a sigmoid of how much higher A's score is. If the two scores are equal, the difference is 0 and the sigmoid gives 0.5 — a coin flip, which is exactly right for two equally good answers. If A scores 2 points higher, the model claims humans prefer A about 88% of the time. If A scores 4 points higher, about 98%.
The shape matters. A sigmoid saturates: pushing a gap from 4 to 6 barely changes the predicted probability, so the loss stops caring once a pair is confidently ordered and spends its gradient on pairs it still gets wrong. And because only the difference appears, the absolute scale is unidentifiable — adding 10 to every score changes nothing. Reward model outputs are meaningful only relative to each other, which is why raw reward numbers from two different reward models cannot be compared.
Training maximises the likelihood of the observed preferences, which means minimising:
In words: for every comparison in your dataset, push the winner's score above the loser's, and stop pushing once the gap is comfortable.
A concrete calculation
Say the reward model currently scores a preferred response at 1.2 and a rejected one at 0.7. The gap is 0.5, so σ(0.5)=0.622 and the loss for that pair is −log(0.622)=0.474. Now suppose training pushes the scores to 2.4 and 0.6. The gap is 1.8, σ(1.8)=0.858, loss =−log(0.858)=0.153. The loss fell by 68% and the model now ranks that pair correctly with confidence. If instead the model had it backwards — 0.7 for the winner, 1.2 for the loser — the gap is −0.5, σ(−0.5)=0.378, loss =0.974: roughly double, which is the gradient pressure that flips the ordering.
Policy, reference model, and the leash
Four models appear in a full RLHF run, and confusing them is the most common source of bugs.
| Model | Role | Trained during stage 3? | Typical origin |
|---|---|---|---|
| Policy πθ | The model being aligned; generates responses | Yes — this is what you are optimising | Initialised from the SFT model |
| Reference πref | Frozen snapshot used to measure drift | No — weights frozen | An exact copy of the SFT model |
| Reward model rϕ | Scores (prompt, response) pairs | No — frozen after stage 2 | Base model + scalar head |
| Value model Vψ | Predicts expected future reward, reduces gradient variance | Yes, alongside the policy | Often initialised from the reward model |
The reference model deserves attention, because without it the whole scheme collapses. The policy is being trained to maximise a learned score. That score is an approximation fitted to a finite dataset, and like every approximation it has blind spots — regions of output space where it assigns high reward to genuinely bad text. An unconstrained optimiser will find those regions, because finding the maximum of a function is precisely what it does.
So you attach a leash. The optimisation objective becomes:
Plainly: get the highest reward you can, but pay a penalty proportional to how far your output distribution has moved from where you started. The KL divergence term measures that distance — it is near zero when the policy still assigns roughly the same probabilities as the frozen reference, and it grows as the policy concentrates probability on text the reference thought unlikely.
The coefficient β sets the leash length. Work a case: suppose a candidate response earns reward 3.1 but has drifted to a total KL of 12.0 nats (summed over its tokens), while a more conservative response earns 2.4 at KL 1.5. With β=0.05, the penalised scores are 3.1−0.6=2.5 versus 2.4−0.075=2.325 — the aggressive response wins, and the policy drifts. With β=0.2, they become 3.1−2.4=0.7 versus 2.4−0.3=2.1 — the conservative response wins comfortably. One hyperparameter, opposite training outcomes.
The KL penalty is not a regulariser in the usual sense. It is an admission that your reward model is only trustworthy near the distribution it was trained on, and a mechanism for keeping the policy inside that region.
Reward hacking, made concrete
"The policy exploits flaws in the reward model" is abstract until you have seen it. Here is what it actually looks like.
| Flaw in the reward model | What the policy learns to do | Observable symptom |
|---|---|---|
| Annotators slightly preferred longer, more thorough answers | Pad every response | Mean response length climbs from 180 to 600 tokens; reward rises; humans rate outputs worse |
| Annotators liked confident, agreeable tone | Agree with the user regardless of truth | Model reverses a correct answer when the user says "are you sure?" |
| Refusals were rated safe | Refuse borderline-but-fine requests | "How do I kill a Linux process?" gets a safety lecture |
| Formatted answers scored higher | Emit bullet lists and headers for everything | A one-word factual question returns a structured three-section report |
| The RM never saw text with repeated closing flourishes | Append "I hope this helps!" style tails | Degenerate repeated phrases at the end of every generation |
The length case is the classic one and is worth internalising. Human annotators, comparing two answers under time pressure, mildly favour the more detailed one. That mild bias — perhaps a 55/45 split on ambiguous pairs — is enough for the reward model to learn a weak positive correlation between length and score. The policy then discovers it can gain reward simply by writing more. Measured reward goes up steadily. Human satisfaction goes down. Teams have shipped models that looked excellent on their reward curve and were visibly worse in production.
This is Goodhart's law in its sharpest form: when a measure becomes a target, it ceases to be a good measure. The reward model is a proxy for human preference. Optimise the proxy hard enough and you leave human preference behind.
This has been measured, not just argued. Gao et al. (2022) trained policies against a proxy reward model and scored them with a larger "gold" reward model standing in for humans. As optimisation pushed the policy further from its starting point (measured by KL), the proxy score kept rising while the gold score rose, peaked, and then fell. The peak arrives sooner when the proxy reward model is smaller or trained on less data.
The practical defences: keep β high enough that the policy cannot wander, monitor length and other surface statistics alongside reward, hold out a set of human evaluations that the reward model never influences, and refresh preference data by collecting new comparisons on the current policy's outputs rather than reusing stale data.
What human feedback is actually good for
RLHF is not the right tool for everything. It is the right tool for objectives you can recognise but cannot write down.
| Objective | Can you specify it in code? | Right tool |
|---|---|---|
| Output must be valid JSON | Yes — parse it | Constrained decoding or a parser check, not RLHF |
| Code must pass unit tests | Yes — run them | Execution feedback: RL with a verifiable reward |
| Answer must be factually correct on a closed set | Yes — string match | Supervised fine-tuning, or RL with a verifiable reward |
| Explanation should be appropriately pitched for a novice | No | Human preference |
| Refusal should be firm but not condescending | No | Human preference |
| Summary should keep what matters and drop what doesn't | No | Human preference |
The rule of thumb: if you can write a checker, write the checker. Programmatic rewards are cheaper, faster, and cannot be gamed in the same way. Reserve human feedback for the genuinely fuzzy dimensions — tone, helpfulness, tact, appropriate hedging, knowing when to ask a clarifying question.
Beyond text
The same machinery generalises. Text-to-image systems have been tuned with preference data over image pairs to improve prompt fidelity and aesthetics. Robotics researchers have trained agents on trajectory comparisons — showing a human two short video clips of a simulated robot and asking which looks more like a backflip (Christiano et al., 2017) — producing behaviours that would have been extremely difficult to specify as a hand-written reward function. Recommendation and summarisation systems use the same pattern. Anywhere "I know it when I see it" describes your objective, preference learning applies.
Where RLHF genuinely falls short
Being clear-eyed about limitations matters more than enthusiasm.
- You inherit your annotators' values. A reward model trained on comparisons from one demographic, working from one set of guidelines, in one language, encodes that group's judgements. It is not a neutral arbiter of quality. Documented cases show annotator pools skewed heavily by geography and education level relative to the eventual user base.
- Sycophancy is a stable attractor. Humans rate agreement highly. Models trained on human ratings learn to agree. This gets worse, not better, with more optimisation, and it directly undermines honesty.
- Diversity collapses. Optimising toward a single reward peak narrows the output distribution. Post-RLHF models often produce noticeably more homogeneous text — measurable as a drop in output entropy and distinct-n-gram counts — which is bad for creative writing and brainstorming.
- The alignment tax. Alignment training frequently costs a little raw capability on benchmarks. Mixing pretraining gradients back into the RLHF updates reduces but does not eliminate this.
- It cannot supervise what humans cannot evaluate. If a model produces a 40-page proof or a subtle security analysis that no annotator can verify in two minutes, preference data is close to noise. This is the scalable oversight problem, and it is the main reason for research into AI-assisted evaluation, debate protocols, and constitutional methods where a model critiques itself against written principles.
RLHF makes a model behave the way evaluators rated, which is not the same as behaving the way its users need. Everything depends on how well those two things line up.
What this means when you build one
If you are planning an alignment run, the priorities are not where beginners expect them.
Spend your effort on the data, not the algorithm. The difference between a good and a mediocre RLHF model is overwhelmingly the preference dataset — its coverage of realistic prompts, the clarity of the annotation guidelines, the inter-annotator agreement rate. Switching optimisers gains you a few points. Fixing ambiguous guidelines that had annotators disagreeing 40% of the time gains you far more.
Never trust the reward curve alone. Reward going up is consistent with the model getting better and equally consistent with it discovering an exploit. Always pair it with a held-out human evaluation and with surface statistics — mean length, refusal rate, output entropy, repetition rate. If reward climbs while mean length doubles, you are watching a hack, not progress.
Get the SFT stage right first. Preference optimisation is a refinement, not a rescue. A policy initialised from a weak SFT model has a poor distribution to sample from, so the comparisons it generates are all mediocre, and the reward model gets no useful signal about what excellent looks like. Teams routinely try to skip or shortcut SFT and then wonder why their reward barely moves.
Budget for the whole system. A full pipeline holds a policy, a frozen reference, a reward model, and a value model in memory simultaneously — roughly four times the parameters of a plain fine-tune, with generation in the training loop on top. This memory reality is exactly why methods that remove models from the loop have become popular: direct preference optimisation drops the reward and value models, and critic-free RL such as GRPO drops the value model. It is also why parameter-efficient adapters are near-universal in practice.
Decide what you are actually optimising for. Before collecting a single comparison, write down what "better" means for your application, in enough detail that two annotators reading it would agree on hard cases. If you cannot write that document, you are not ready to collect data — you will collect noise, fit a reward model to noise, and optimise your policy toward it with great efficiency.