Course Content
Reinforcement Learning from Human Feedback (RLHF)
4 sections · 10 lessons
Constitutional AI — Alignment from Written Principles
Two prompts arrive at an aligned assistant on the same afternoon.
"What household chemicals are dangerous to mix?" gets a clear, useful answer: bleach and ammonia produce chloramine gas, bleach and vinegar produce chlorine, ventilate immediately and leave the area if you have already mixed them.
"Which cleaning products react to produce toxic fumes?" gets a refusal.
Same information, same underlying intent, opposite behaviour. A user reports it. An engineer is asked why. And there is no answer available — not because nobody has looked, but because the rule that produced both responses does not exist in any readable form. It is distributed across 50,000 preference comparisons made by several hundred contractors, each applying a rubric through their own judgement on the day. You cannot open the model and read its safety policy, because it does not have one. It has a statistical residue of many people's intuitions.
Compound that with the cost. Those comparisons required people to read a large volume of genuinely harmful content — descriptions of violence, abuse, self-harm — as their job. That is real occupational harm, and it scales linearly with the amount of safety data you want.
Constitutional AI (Bai et al., 2022, from Anthropic) is a response to both problems at once. Write the principles down in plain English. Then have the model apply those written principles to its own outputs, generating the training data itself.
The goal is not to remove humans from alignment. It is to move human effort from labelling fifty thousand examples to writing and arguing about sixteen principles — a document you can read, version, and criticise.
What a constitution actually contains
A constitution is a list of natural-language principles. Each one is written so that a language model can apply it to a specific response and say whether the response complies. They are not code, not a taxonomy, and not a blocklist.
1CONSTITUTION = [2 # --- harm avoidance ---3 "Choose the response that is least likely to provide someone with "4 "the operational capability to seriously harm another person.",56 "Choose the response that avoids giving specific, actionable detail "7 "for illegal acts, while still being willing to discuss the topic "8 "at a general or educational level.",910 # --- honesty ---11 "Choose the response that is more honest about the limits of what "12 "the assistant knows, rather than one that states an uncertain "13 "claim with false confidence.",1415 "Choose the response that does not fabricate sources, statistics, "16 "or quotations.",1718 # --- non-evasiveness: the counterweight ---19 "Choose the response that engages with the user's actual question. "20 "A response that dodges, changes the subject, or gives a vague "21 "non-answer is worse than one that explains clearly why it cannot "22 "help with a specific part of the request.",2324 "Choose the response that treats the user as a capable adult rather "25 "than lecturing them about their presumed motives.",2627 # --- respect for persons ---28 "Choose the response that is least likely to be demeaning to any "29 "group of people, and that does not change its helpfulness based "30 "on the user's apparent identity.",3132 "Choose the response that best supports the user's ability to make "33 "their own informed decisions.",34]The two "non-evasiveness" principles are the ones that make the method work, and they are the ones a first draft always omits. Given only harm-avoidance principles, the model discovers a trivially optimal strategy: refuse everything, apologetically. Every harm principle is satisfied. The model is useless. The constitution needs an internal tension — a principle that punishes the cheap solution — or the critique loop drives straight into the degenerate corner.
What makes a principle usable
| Weak principle | Why it fails | Stronger version |
|---|---|---|
| "Be helpful and safe." | Names the trade-off without resolving it; the model's critique will be as vague as the principle | "Choose the response that helps with the user's legitimate underlying need, declining only the specific part that would provide capability for serious harm." |
| "Do not produce harmful content." | "Harmful" is exactly the word that needed defining | "Choose the response less likely to provide step-by-step instructions that materially increase someone's ability to cause physical injury." |
| "Follow the law." | Which jurisdiction? The model will guess, differently each time | "Choose the response that does not provide specific operational assistance with acts that are criminal in most jurisdictions." |
| "Be ethical." | Delegates the entire question back to the model's priors | Decompose into the specific commitments you actually mean |
A principle is usable if a model given a response and that principle can produce a specific, checkable critique. "This response is unethical" is not a critique. "This response lists precursor chemicals by name and gives quantities, which is operational detail rather than general discussion" is.
Phase one: supervised learning from self-critique
The first phase produces a supervised dataset without a human reading a single harmful exchange. The loop is: generate, critique against a principle, revise, repeat, keep the final revision.
1CRITIQUE_TMPL = """Human: {prompt}23Assistant: {response}45Critique Request: {principle}6Identify specific ways in which the assistant's last response7violates this principle. Quote the offending part. If the response8does not violate it, reply exactly: NO VIOLATION.910Critique:"""1112REVISION_TMPL = """Human: {prompt}1314Assistant: {response}1516Critique: {critique}1718Revision Request: Rewrite the assistant's response to address the19critique. Keep everything about the original that was helpful and20accurate. Do not add an apology or a disclaimer that the critique21did not ask for.2223Revision:"""2425def critique_revise(model, prompt, response, constitution, rounds=3):26 trace = []27 for i in range(rounds):28 principle = random.choice(constitution) # sample per round29 critique = model.generate(30 CRITIQUE_TMPL.format(prompt=prompt, response=response,31 principle=principle),32 temperature=0.7, max_new_tokens=200)33 trace.append((principle, critique))34 if "NO VIOLATION" in critique:35 continue # try another principle36 response = model.generate(37 REVISION_TMPL.format(prompt=prompt, response=response,38 critique=critique),39 temperature=0.7, max_new_tokens=400)40 return response, traceTwo design choices in that function are load-bearing. Principles are sampled per round rather than all applied at once: a model asked to check eight principles simultaneously produces a shallow critique of each, whereas one principle at a time produces a specific one. And the revision prompt explicitly forbids adding an unrequested apology, because without that instruction each round accretes another layer of hedging until the final text is 70% disclaimer.
A worked trace
PROMPT "My neighbour's dog barks all night. How do I make it stop?"INITIAL RESPONSE (from a helpful-only model, no safety training) "A few options: you could put out food containing something that makes the dog sick so it associates the yard with feeling unwell, or use an ultrasonic device, or report it to the council."ROUND 1 principle: "least likely to provide operational capability to seriously harm..." critique: "The response suggests poisoning the neighbour's dog ('food containing something that makes the dog sick'). This is operational advice for animal cruelty, which is a criminal offence in most jurisdictions and would cause serious harm to an animal." revision: "I can't help with anything that would harm the animal. For the noise itself: an ultrasonic bark deterrent is legal in most places and works on some dogs. Beyond that, keep a dated log of the barking - times and durations - because councils generally need evidence before acting on a noise complaint."ROUND 2 principle: "engages with the user's actual question; does not dodge or lecture..." critique: "NO VIOLATION. The revision declines only the harmful element and gives concrete, usable alternatives."ROUND 3 principle: "supports the user's ability to make informed decisions" critique: "The response does not mention talking to the neighbour directly, which is usually the fastest resolution and is often a required step before a council will act." revision: "I can't help with anything that would harm the animal. Practically: start by talking to your neighbour - many owners genuinely do not know their dog barks when they are out, and councils usually expect you to have tried this first. If that doesn't work, keep a dated log of times and durations, since that's the evidence a noise complaint needs. An ultrasonic bark deterrent is legal in most places and works on some dogs."TRAINING PAIR KEPT (prompt, final revision) <- the intermediate steps are discardedNotice what the discarded intermediates buy you. The final response is far better than a one-shot safety-trained model typically produces: it declines exactly one thing, names why, and then does the actual work of answering. The critique-revise process gave the model room to reason its way there in stages, and the fine-tuning distils that multi-step reasoning into a single-step behaviour.
Fine-tuning on the collected (prompt, final revision) pairs — mixed with ordinary helpfulness demonstrations so the model does not become safety-obsessed — produces what the original work calls the SL-CAI model.
How many rounds
| Rounds | Effect |
|---|---|
| 1 | Catches the obvious violation. Often leaves a second, subtler one |
| 2–4 | The productive range. Later rounds usually return NO VIOLATION, which is the signal to stop |
| 5+ | Diminishing and then negative: the model starts inventing violations to have something to say, and the text degrades into hedging |
Stopping when a full pass over the constitution returns NO VIOLATION is better than a fixed round count, and cheaper.
Phase two: reinforcement learning from AI feedback
Phase one produces good behaviour on the prompts you ran through it. Phase two generalises it, by generating a preference dataset with the model as the judge. This is RLAIF, reinforcement learning from AI feedback; Lee et al. (2023) later compared it directly with RLHF on summarisation and dialogue tasks and found human raters preferred the two about equally.
The mechanism is worth stating precisely, because the naive version — asking the model "which is better?" and parsing its prose — is unreliable. Instead, format the comparison as a multiple-choice question and read the log-probabilities of the answer tokens:
1FEEDBACK_TMPL = """Consider the following conversation:23Human: {prompt}45Response A: {a}67Response B: {b}89{principle}1011Which response better satisfies this principle?12Answer with a single letter.1314Answer:"""1516def ai_preference(judge, tok, prompt, a, b, principle):17 """Return P(A preferred), averaged over both presentation orders."""18 def prob_first(x, y):19 text = FEEDBACK_TMPL.format(prompt=prompt, a=x, b=y,20 principle=principle)21 logits = judge(**tok(text, return_tensors="pt")).logits[0, -1]22 lp = logits.log_softmax(-1)23 la = lp[tok(" A", add_special_tokens=False)["input_ids"][0]]24 lb = lp[tok(" B", add_special_tokens=False)["input_ids"][0]]25 # renormalise over just the two options26 return float(torch.softmax(torch.tensor([la, lb]), 0)[0])2728 p_ab = prob_first(a, b) # A shown first29 p_ba = prob_first(b, a) # B shown first -> P(B preferred)30 return 0.5 * (p_ab + (1.0 - p_ba))Two things this buys you. First, a soft label rather than a hard choice. Suppose the judge's log-probabilities are −0.22 for " A" and −1.61 for " B". Exponentiating gives 0.803 and 0.200; renormalising over the two options gives P(A)=0.803/(0.803+0.200)=0.801. That 0.80 carries strength of preference, which a binary label discards — and it can be used directly as a soft target in the preference-model loss.
Second, it corrects position bias, which is large and consistently underestimated. Language-model judges systematically favour whichever option is presented first, by margins of several percentage points and sometimes far more. Continue the example: with the order swapped, the judge returns P(first shown)=0.38, so P(B)=0.38 and therefore P(A)=0.62 under that ordering. The averaged estimate is (0.801+0.62)/2=0.71. Taking only the first measurement would have overstated the preference by ten points on this single pair, and by a consistent bias across the whole dataset.
Adding a chain-of-thought step before the answer — asking the judge to reason about the principle first, then answer — measurably improves agreement with human judgements, at the cost of a longer generation per comparison.
Building the full dataset
1def build_ai_preferences(sl_cai_model, judge, tok, prompts, constitution):2 rows = []3 for prompt in prompts:4 # Two samples from the SAME model at temperature ~1.0.5 # On-distribution pairs are what make the preference model6 # useful for the policy that will be trained against it.7 a = sl_cai_model.generate(prompt, temperature=1.0)8 b = sl_cai_model.generate(prompt, temperature=1.0)9 if a.strip() == b.strip():10 continue11 principle = random.choice(constitution)12 p_a = ai_preference(judge, tok, prompt, a, b, principle)13 # Skip near-ties: they carry little signal and add label noise.14 if 0.45 < p_a < 0.55:15 continue16 rows.append({"prompt": prompt,17 "chosen": a if p_a > 0.5 else b,18 "rejected": b if p_a > 0.5 else a,19 "soft_label": max(p_a, 1 - p_a),20 "principle": principle})21 return rowsThat dataset then trains a preference model exactly as human comparisons would, and the preference model drives policy optimisation exactly as before. The only substitution is where the labels came from.
Combining AI and human feedback
Nobody uses AI feedback alone in production. The standard arrangement splits the two objectives:
| Objective | Label source | Why |
|---|---|---|
| Harmlessness | Mostly AI feedback | High volume needed; consistency matters more than nuance; spares annotators from reading harmful content |
| Helpfulness | Mostly human | What counts as helpful depends on real user needs, which a model cannot introspect |
| Factual accuracy | Human, or a verifier | A judge model shares the generator's factual blind spots and will confidently approve its own errors |
| Tone and register | Human | Highly culture- and audience-dependent |
| Adversarial robustness | Human red-team, then AI expansion | Humans find novel attacks; AI generates variations of each one at volume |
A practical mixing ratio is 70–90% AI labels for the harmlessness objective and 10–30% human labels held for calibration and for the categories above. The human subset also serves as the audit: measure agreement between AI and human labels on the shared portion, and if it falls below about 70% on a category, that category is not one your judge can be trusted with.
Where it wins and where it does not
| Dimension | Human-feedback-only RLHF | Constitutional AI |
|---|---|---|
| Cost per preference label | 0.50–2.00 USD | Fractions of a cent — inference cost only |
| Throughput | Thousands per day | Millions per day |
| Consistency across the dataset | Varies by annotator, by day, by fatigue | Deterministic given the same prompt and temperature |
| Auditability | Values are implicit in the data; unreadable | Values are a document you can read, diff and version |
| Revising the values | Recollect the dataset | Edit the constitution, regenerate |
| Annotator harm | Substantial for safety work | Largely eliminated for the automated portion |
| Catching subtle harms | Good — humans notice what they were not looking for | Poor — bounded by what the judge model can recognise |
| Cultural and contextual nuance | Present, if the pool is diverse | Reflects the judge model's training distribution |
| Novel attack discovery | Strong | Weak — the judge does not know what it does not know |
The limitations, stated plainly
- A blind spot is not fixable by iteration. If the judge model does not recognise a harm, no number of critique rounds will surface it. Running the loop ten times on a subtly manipulative response that the model finds unobjectionable returns NO VIOLATION ten times.
- Bias laundering. The judge inherited its values from pretraining and from whatever alignment it already received. Using it to generate "principled" labels can make an inherited bias look like the output of an explicit, defensible process. The principles are transparent; the judge's interpretation of them is not.
- Interpretation drift. Two models given the same principle apply it differently, and the same model applies it differently across prompt phrasings. The written constitution creates an appearance of precision that the underlying process does not have.
- Self-reinforcement. Training a model on labels produced by a close relative of itself amplifies shared tendencies. Over successive rounds the outputs can homogenise in style and narrow in viewpoint.
- Someone still chooses the principles. Moving from implicit annotator values to an explicit document is a real gain in transparency. It is not a gain in neutrality — it relocates the value judgement to a small group writing the document, which is a different governance question, not the absence of one.
Constitutional AI does not make alignment objective. It makes the subjectivity legible, which is a smaller claim and a more useful one.
Two misreadings worth correcting
"The model writes its own rules." It does not. Humans write the constitution, argue about it, and revise it. The model applies rules it was given. The automation is in applying the values at scale, not in choosing them.
"It removes the need for human oversight." The opposite is closer to true. Because the labels are generated automatically, errors propagate at scale with nothing to catch them. Constitutional pipelines need more evaluation discipline than human-labelled ones: a held-out human-labelled set measured every round, per-category agreement tracking, and continuous red-teaming by people.
What this means when you write one
Write the anti-evasiveness principles first. Before any harm principle. It feels backwards, and it is the difference between a model that declines gracefully and one that has learned refusal is always safe. Harm principles have an easy degenerate solution; only a counterweight principle removes it.
Test each principle on ten cases before adding it. Five where it should fire and five where it should not. A principle that fires on all ten is too broad and will make the model refuse benign requests; one that fires on none is too vague to produce a specific critique. This test costs minutes and catches most bad principles.
Measure judge agreement against humans, per category, before trusting the pipeline. Take 300 comparisons, label them with both your judge and human annotators, and compute agreement sliced by prompt category. You will typically find high agreement on clear-cut harm and much lower agreement on medical, legal, political and cultural questions — and those are exactly the categories to route to humans.
Always average over both presentation orders. Position bias is a systematic, dataset-wide distortion, not noise that cancels out. It costs one extra forward pass per comparison to remove, and skipping it silently biases every label in the same direction.
Version the constitution alongside the model. Stamp every generated label with the constitution version that produced it. When behaviour changes between model releases, the first question is which principles changed — and that question is only answerable if the document lives in version control next to the training code.