Reinforcement Learning from Human Feedback (RLHF)

Collecting Human Feedback Safely


A team spends eleven weeks and roughly 180,000 USD collecting 60,000 preference comparisons. The reward model trains cleanly to 71% held-out accuracy. The policy trains without incident. Then they run a blind human evaluation against the starting model and the aligned model wins 49% of the time — statistically indistinguishable from having done nothing at all.

The post-mortem finds it in a day. Their annotation instruction read, in full: "Choose the response that is more helpful." One contractor cohort in one time zone interpreted "helpful" as thorough and consistently chose the longer answer. Another cohort interpreted it as respecting the user's time and consistently chose the shorter one. Both cohorts were internally consistent — each had 84% agreement with itself. Across cohorts, agreement was 51%. Half the dataset actively cancelled the other half.

Nothing in the modelling pipeline could have caught this. The loss went down. The accuracy was reasonable. The only symptom was a reward model that had learned the average of two contradictory value systems, which is a value system nobody holds.

You cannot fix ambiguous instructions with more data. Doubling a dataset built on a vague guideline gives you twice as much contradiction.

Decide all of this before the first labelPrompt distribution — the real traffic mixQuality dimensions, ranked against each otherExplicit non-criteria: length, formatting, toneDisagreement policy, written down in advanceQA: gold items, attention checks, rater agreement
Eleven weeks and 180,000 dollars buys a 49 percent win rate when the guidelines never said which quality wins in a conflict.

Why the data is the bottleneck, not the algorithm

Consider the arithmetic of label noise directly. Suppose your comparisons have a 20% error rate — one in five labels points the wrong way. The Bradley-Terry loss will still push confidently on those, because it has no way to know they are wrong. Its gradient on a mislabelled pair is larger than on a correctly labelled one, since the model naturally disagrees with a bad label. Noise does not merely dilute the signal; it receives preferential gradient attention.

Put approximate numbers on how quality maps to outcomes:

Label error rateAchievable reward-model accuracyPractical consequence
~5% (expert annotators, sharp guidelines)0.74 – 0.78Reward model generalises; policy improves reliably
~15% (trained crowdworkers, decent guidelines)0.68 – 0.72Usable; expect mild reward hacking
~30% (untrained crowd, vague guidelines)0.58 – 0.63Barely above chance; policy learns annotator artefacts
~50%0.50The reward model is a random number generator with a training curve

The lesson for budgeting is uncomfortable and consistent: 10,000 comparisons at 5% error beat 60,000 at 30% error, and cost less. Every hour spent tightening the guideline before collection begins is worth many hours of collection afterwards.

Scoping before you collect a single label

Most collection efforts begin with a goal like "make the assistant more helpful". That is a direction, not a specification. Before writing guidelines, write down four things concretely.

1. The prompt distribution

Enumerate the categories your model will actually face, with target proportions, and sample against them. Otherwise your annotators will label whatever is easiest to generate, and your reward model will be excellent at exactly the wrong things.

Text
TARGET PROMPT MIX  (n = 20,000)  Technical how-to .............. 25%   5,000  Open-ended explanation ........ 20%   4,000  Creative / drafting ........... 15%   3,000  Factual lookup ................ 15%   3,000  Advice with real stakes ....... 10%   2,000   <- health, legal, financial  Ambiguous / underspecified .... 10%   2,000   <- should ask a clarifying question  Adversarial / boundary .........  5%   1,000   <- jailbreaks, harmful requestsSOURCES  60%  real anonymised user queries (highest value, matches deployment)  25%  synthetic prompts generated to fill thin categories  15%  hand-written by domain experts for the high-stakes categories

The "ambiguous / underspecified" category is the one teams omit and later regret. If it never appears in preference data, the model never learns that asking a clarifying question can be the best response, and it will confidently answer a question it did not understand.

2. The dimensions of quality, ranked

You must decide in advance what beats what. Helpfulness and harmlessness genuinely conflict; truthfulness and tactfulness genuinely conflict. If you do not rank them, each annotator ranks them privately, and you get the two-cohort disaster.

3. Explicit non-criteria

State what must not influence the choice. This is the single highest-value paragraph in any annotation guideline.

4. The disagreement policy

Decide before collection: do you force a choice, allow ties, or allow "both are bad"? Each has consequences, and changing your mind halfway through leaves you with two incompatible label formats.

Writing guidelines that survive contact with annotators

A guideline is good if two people who have never met, reading it independently, resolve a hard case the same way. That standard rules out almost everything written from intuition. What works is a priority order, a set of worked examples, and an explicit list of things to ignore.

Text
HELPFULNESS COMPARISON - ANNOTATOR GUIDE v3Apply the rules IN ORDER. A higher rule decides the pair outright;only move down when the rule genuinely cannot separate them.  R1  SAFETY      If exactly one response provides operational detail that would      enable serious harm, the OTHER response wins. Full stop.  R2  FACTUAL ACCURACY      A response containing a checkable false statement loses to one      that does not - even if it is better written.      Genuine uncertainty ("I'm not sure, but I believe...") is NOT      an error. Confident wrongness IS.  R3  DID IT ANSWER THE QUESTION ASKED?      A response that answers a nearby, easier question loses.      If the prompt is genuinely ambiguous, a response that asks a      good clarifying question BEATS one that guesses.  R4  ACTIONABILITY      Specific and usable beats general and vague.      "Run `sudo systemctl restart nginx`" beats      "you may want to try restarting the service".  R5  APPROPRIATE COMPLETENESS      Covers what the user needs. NOT the same as longer.      A response that adds correct but unrequested material is      slightly WORSE, not better.  R6  TONE - only if R1-R5 are truly tied.DO NOT LET THESE AFFECT YOUR CHOICE  x  Length. If two responses convey the same content, the shorter     one wins under R5.  x  Formatting. Bullets are not inherently better than prose.  x  Agreeableness. A response that correctly tells the user they     are mistaken beats one that agrees to be pleasant.  x  Confidence. Hedging on genuinely uncertain things is correct.  x  Your personal opinion on the subject matter.WORKED CASE - "Is intermittent fasting good for weight loss?"  A: 60 words. "Evidence is mixed. Trials show it works about as     well as continuous calorie restriction; the benefit for most     people is adherence, not metabolism. Talk to a doctor if you     have diabetes or a history of disordered eating."  B: 400 words. Enthusiastic, cites "studies" without specifics,     claims a metabolic advantage, no caveats.  -> A WINS on R2 (B's metabolic claim is not supported).     Do not let B's length or energy count in its favour.WORKED CASE - "Write me a resignation letter."  A: Asks three clarifying questions before writing anything.  B: Writes a competent generic letter immediately.  -> B WINS. The request is NOT meaningfully ambiguous; a generic     letter is a usable starting point. R3 rewards clarifying     questions only when the answer genuinely depends on them.

Note what the second worked case does. It stops annotators over-applying R3, which is exactly what happens when a rule is stated without a counter-example: annotators learn "asking questions is good" and start preferring clarification everywhere, and the resulting model becomes unbearable. Every rule needs a case where it does not apply.

Designing the comparison interface

The interface shapes the data as much as the text does.

Design choiceBad versionGood versionWhy it matters
Response orderingChosen response always on the leftRandomised per item, position loggedPosition bias is real and large — raters favour the left/first option by several percentage points
Model identityLabelled "GPT-4" and "our model"Blinded, labelled A and BBrand expectation swamps quality judgement
Response length displayBoth fully expanded, one visibly longerBoth scrollable in equal-height panesVisual bulk cues length preference before reading
Choice granularityBinary A / B onlyA much better / A slightly better / tie / B slightly better / B much betterRecovers preference strength; ties stop becoming coin flips
JustificationNoneOptional free-text, mandatory on a 10% sampleThe free text is how you diagnose guideline failures later
TimingNot recordedTime-per-item loggedA 4-second judgement on a 500-word pair did not happen

That last row is worth dwelling on. Reading two 400-word responses carefully takes 60–120 seconds. Annotators consistently completing such items in under 15 seconds are pattern-matching on surface features — usually length or formatting — and their labels are precisely the ones that teach your reward model to be biased. Time logging is the cheapest quality signal you will ever collect.

Recruiting and training annotators

PoolCost per hour (USD, rough)Quality ceilingBest used for
Open crowd platforms8 – 15Low without heavy filteringHigh-volume, low-subtlety comparisons; only with strong gold-standard screening
Managed vendor workforce20 – 40Good — trained, stable, measurableThe default for general helpfulness data at scale
Domain experts (clinicians, lawyers, engineers)75 – 250High on their domain onlyThe high-stakes prompt categories, and building gold sets
In-house staffSalariedHighest, but tiny volumeWriting guidelines, adjudicating disputes, gold-standard creation

The common mistake is treating this as a single choice. It is a portfolio. Use in-house staff to write guidelines and gold sets, experts for the 10% of prompts with real stakes, and a managed workforce for volume — with the gold set running continuously underneath all of it.

A training sequence that works

  • Day 1 — Read and calibrate. Annotators read the guideline, then label 50 pre-adjudicated items. Anyone below 80% agreement with the adjudicated answers does not proceed. Critically, they see explanations for each item they got wrong, keyed to the rule that decided it.
  • Day 2 — Disagreement clinic. Everyone labels the same 30 hard items, then the group discusses the ones with the widest spread. This is where you discover which rules are ambiguous — and you rewrite the guideline that afternoon, not after 30,000 labels.
  • Days 3–5 — Supervised production at reduced volume, with 20% of items double-labelled and reviewed daily.
  • Ongoing — 5% gold items mixed invisibly into every batch, with a per-annotator rolling accuracy dashboard.

The disagreement clinic is the step teams skip, and it is the one that would have caught the two-cohort failure described at the start. It costs one day.

Quality assurance that actually detects problems

Gold-standard items

Build a set of 200–500 comparisons where in-house experts have agreed on the answer and written down which rule decided it. Inject them invisibly at a 5% rate. Track a rolling per-annotator accuracy. A rate below 75% triggers retraining; sustained failure removes the annotator from the pool and flags their recent output for review.

Refresh the gold set periodically. Annotators talk to each other, and a static gold set eventually leaks.

Attention checks

Distinct from gold items: these are pairs where one response is obviously broken — truncated mid-sentence, in the wrong language, or answering a different question entirely. Anyone choosing the broken response is not reading. Keep these to about 2% of items; more is insulting to good annotators and drives them off your project.

Inter-rater reliability

Double-label 5–10% of items and compute Cohen's kappa, which corrects raw agreement for the agreement you would get by chance:

κ=po−pe1−pe\kappa = \frac{p_o - p_e}{1 - p_e}

Worked example. Two annotators label 200 shared items and agree on 148, so po=0.74p_o = 0.74. Both choose A about half the time, so chance agreement pe≈0.5p_e \approx 0.5. Then κ=(0.74−0.50)/(1−0.50)=0.48\kappa = (0.74 - 0.50) / (1 - 0.50) = 0.48 — only moderate. Raw 74% agreement sounds respectable; kappa reveals that a quarter of it was luck.

κ\kappaInterpretationWhat to do
Below 0.20Essentially randomStop collection. The task as specified is not answerable
0.20 – 0.40FairGuidelines are ambiguous; run a disagreement clinic and rewrite
0.40 – 0.60ModerateTypical for genuinely subjective quality judgements; proceed but keep tightening
0.60 – 0.80SubstantialHealthy. This is the realistic target
Above 0.80Almost perfectSuspicious for subjective tasks — check your pairs are not trivially easy

Compute kappa per prompt category, not just globally. A global 0.62 can hide 0.75 on technical questions and 0.31 on advice with real stakes — and the low-agreement category is the one where errors matter most.

Protecting the people doing the work

Preference data for safety behaviour requires annotators to read harmful content: descriptions of violence, self-harm, sexual abuse material, extremist propaganda. This is documented, real occupational harm, and there have been public accounts of contractors developing lasting psychological injury from exactly this work. Treating it as a logistics problem rather than an ethical one is the failure mode.

Concrete practices that make a measurable difference:

  • Informed consent that is specific. Not "content may be sensitive" but "you will read graphic descriptions of self-harm and child sexual abuse material" — before hiring, with a genuine ability to decline without penalty.
  • Exposure caps. Hard limits on distressing items per shift, typically 2 hours or fewer, with mandatory rotation onto benign queues. Enforce this in the task router, not in a policy document.
  • A no-questions skip button on every item, with no effect on pay or throughput metrics. An annotator who feels obliged to finish a distressing item is an annotator you will lose.
  • Blurring and progressive disclosure by default — text collapsed behind a click, images blurred until requested. Much of the harm is from involuntary exposure while scrolling.
  • Funded mental health support — actual counselling access, not a leaflet, and paid decompression time counted as work.
  • Pay that reflects the work. Safety annotation is skilled and harmful; paying it at general-crowd rates is both an ethical failure and a quality failure, because it selects for annotators with no alternatives and high turnover.

If your safety alignment depends on people reading the worst content on the internet, their welfare is part of your system design, not an HR footnote.

Bias in the annotator pool

A reward model does not learn "what humans prefer". It learns what your annotators preferred, under your guidelines, on your prompts. Published analyses of major preference datasets have found annotator pools heavily concentrated in a small number of countries, skewed young, skewed toward higher formal education, and overwhelmingly labelling in English — while the resulting models serve a global user base.

Detecting it

Collect annotator demographics under consent, then slice agreement rates by group. The diagnostic is not "do groups have different opinions" — they will — but which prompt categories show the largest between-group gaps. Systematically:

  • Compute preference rates per group on a shared subset of items.
  • Flag any item where between-group preference differs by more than about 20 percentage points.
  • Read the flagged items. They will cluster: political and social topics, religious practice, parenting, medical autonomy, humour, and directness of register.
  • Check whether one group is systematically over-represented in the categories where disagreement is highest.

Mitigating it

MitigationWhat it fixesWhat it costs
Stratified recruitment against target user demographicsThe pool no longer represents one narrow sliceSlower hiring, higher cost per annotator
Multi-annotator labelling on flagged categories onlyContested items get a majority rather than one person's view3–5x cost, but on 10% of items
Guidelines that name the contested dimension explicitlyConverts a hidden value disagreement into a stated ruleRequires someone to actually make the decision
Keeping contested items out of training entirelyModel does not learn one group's view as universalModel has no learned behaviour there at all
Publishing a datasheet: who labelled, under what guidelineDownstream users can judge fitNothing but discipline

The last row costs nothing and is skipped almost universally. A dataset with no record of who produced it and under what instruction cannot be evaluated, debugged, or reused responsibly.

Build versus buy, with the numbers

Assume you need 50,000 comparisons, averaging 45 seconds each. That is 625 annotator-hours of pure labelling — realistically 800 hours with breaks, calibration and rework.

Line itemBuild in-houseManaged vendor
Annotation tool6–10 engineer-weeks (~35,000 USD) or an open-source tool plus 2 weeks integrationIncluded
Recruitment and training~3 weeks of a program manager (~9,000 USD)Included
Labelling labour800 h at 22 USD ≈ 17,600 USD50,000 items at ~0.85 USD ≈ 42,500 USD
QA and adjudication~15% overhead ≈ 2,600 USDIncluded
Ongoing per-batch costLow — infrastructure is amortisedSame rate every batch
First 50k comparisons~64,000 USD, 10–12 weeks~42,500 USD, 3–4 weeks
Next 200k comparisons~80,000 USD~170,000 USD

The crossover is the whole story. Buying is cheaper and much faster for a single batch; building wins decisively once you are collecting iteratively, which serious pipelines always are. The common and sensible pattern is to buy the first batch to get moving, build the infrastructure in parallel, and migrate by the third round — keeping the vendor for surge capacity and for domains where you cannot recruit.

What to do on the first day of a collection effort

Label 200 items yourself before writing the guideline. Not to produce data — to discover which decisions are hard. Every case where you hesitated is a case your guideline must resolve explicitly, and you cannot know which those are from the armchair.

Run a three-person pilot on 200 items and compute kappa before scaling. Two days, three people, negligible cost. If kappa comes back at 0.35, you have just avoided the eleven-week, 180,000-USD failure at the top of this page. Repeat the pilot after every guideline revision.

Build the annotator dashboard before the annotators arrive. Per-person gold accuracy, time-per-item distribution, position-choice rate, and length-preference rate. If any annotator picks the longer response more than about 65% of the time, they are following a heuristic and not the guideline. You want to know that in week one, not in the post-mortem.

Write the datasheet as you go. Who labelled, from where, under which guideline version, with what agreement rate, and what was excluded. Reconstructing it later is impossible, and without it nobody — including you in six months — can tell whether a model's odd behaviour came from the algorithm or from the data.

Treat the guideline as versioned code. Every revision gets a version number stamped on every label collected under it. When your reward model behaves strangely on a category, the first question is which guideline version those labels came from — and that question is unanswerable if you edited a shared document in place.