Course Content
Reinforcement Learning from Human Feedback (RLHF)
4 sections · 10 lessons
Open-Source RLHF Frameworks — TRL, Axolotl, TRLX
An engineer decides to write PPO from scratch. It is not unreasonable — the clipped objective is about twenty lines, GAE is another fifteen, and the paper is clear. Three days later it runs correctly on a 125M model on one GPU. The reward goes up. Everything works.
Then it has to run on a 7B model, and the next three weeks look like this. It runs out of memory, so ZeRO stage 3 shards the optimiser states — and now model.generate() silently produces garbage, because the parameters are sharded across ranks and generation needs them gathered. Gradient checkpointing halves activation memory but conflicts with the value head's forward hook. Mixed precision makes the KL estimate unstable at bf16. Multi-GPU rollouts need the batch resharded between the generation phase and the optimisation phase. The tokeniser pads right for training and needs to pad left for generation, and getting that wrong shifts every log-probability by one position without raising an error.
None of that is PPO. All of it is distributed-systems plumbing, and every one of those problems has already been solved, argued about, and fixed in a library.
The value of an RLHF framework is almost never the loss function. It is that someone has already reconciled sharded parameters with autoregressive generation, and you did not have to.
The landscape
| Tool | Interface | Covers | Scale ceiling | Status |
|---|---|---|---|---|
| TRL (Hugging Face) | Python API | SFT, reward models, DPO and variants, KTO, GRPO, RLOO; more in trl.experimental. No PPO in 1.x | Single node to modest multi-node | Actively developed; the default for most work |
| Axolotl | YAML config + CLI | SFT, LoRA/QLoRA, DPO, ORPO, KTO, GRPO | Multi-GPU, multi-node via DeepSpeed/FSDP | Active; the fastest path from data to a trained adapter |
| TRLX (CarperAI) | Python API | PPO, ILQL | Was the multi-node option of its era | Effectively unmaintained — historically important, not a choice for new work |
| LLaMA-Factory | YAML + web UI | SFT, PPO, DPO, KTO, many model families | Multi-GPU | Active; broadest model support |
| OpenRLHF | Python + shell scripts | PPO, REINFORCE++, GRPO, RLOO, plus SFT, reward models and DPO | Large multi-node; Ray-based, vLLM for rollouts | Active; built for the scale where rollout generation dominates |
| veRL (ByteDance) | Python + Hydra config | PPO, GRPO, DAPO, RLOO, REINFORCE++ and more | Large multi-node; FSDP or Megatron for training, vLLM or SGLang for rollouts | Active; widely used for RL on reasoning models |
| Unsloth | Python API | SFT, DPO, GRPO with custom fused kernels | Single GPU | Active; large speed and memory wins on one card |
Two of these are worth knowing properly, because between them they cover most real work: TRL when you need programmatic control, Axolotl when you need a reproducible pipeline fast. For PPO with a critic, or RL at a scale where generation dominates, use OpenRLHF or veRL; section 2 shows an OpenRLHF launch.
A note on TRLX
TRLX mattered. When RLHF first became reproducible outside large labs, it was the library that made distributed PPO on multi-billion-parameter models practical, and it introduced the pattern of a reward function passed in as a plain Python callable:
1# The TRLX idea, which everything since has borrowed:2# training is driven by an arbitrary function from text to float.3def reward_fn(samples, prompts=None, outputs=None):4 return [float(score(rm, rm_tok, s)) for s in samples]56# trlx.train(reward_fn=reward_fn, prompts=prompts, config=cfg)That callable interface is the right abstraction and it survives in every modern library. The project itself, however, has not kept pace with current model architectures and training stacks. Treat it as background: if you inherit a TRLX codebase, understand it; do not start one.
TRL in practice
1pip install "trl>=1.0,<2" "transformers>=5" peft accelerate \2 datasets bitsandbytes34# Pin these. The code below was checked against trl 1.13. TRL's5# trainers have changed across versions more than once (PPOTrainer6# was removed; max_prompt_length left DPOConfig), and code from an7# older tutorial frequently fails on a renamed argument.8pip freeze | grep -E "^(trl|transformers|peft|accelerate)=="Stage one: supervised fine-tuning
1from trl import SFTTrainer, SFTConfig2from peft import LoraConfig34cfg = SFTConfig(5 output_dir="sft-out",6 max_length=2048,7 packing=True, # concatenate short examples to fill8 # the context - large throughput win,9 # but see the warning below10 num_train_epochs=2,11 per_device_train_batch_size=4,12 gradient_accumulation_steps=4,13 learning_rate=2e-4, # high because LoRA; 2e-5 for full FT14 bf16=True,15)1617trainer = SFTTrainer(18 model="meta-llama/Llama-3.2-1B",19 args=cfg,20 train_dataset=ds, # column: "text" or "messages"21 peft_config=LoraConfig(r=16, lora_alpha=32, lora_dropout=0.05,22 target_modules=["q_proj", "k_proj",23 "v_proj", "o_proj"]),24)25trainer.train()packing=True deserves the warning. It concatenates multiple short examples into one sequence to avoid wasting context on padding, which can double throughput on datasets of short turns. It also means examples bleed across boundaries unless the implementation masks attention between them. For instruction data with short responses this occasionally teaches the model to continue past its own end-of-turn token into the next example. If your fine-tuned model starts inventing a follow-up user turn after answering, packing is the first thing to switch off.
Stage two: the reward model
1from trl import RewardTrainer, RewardConfig23cfg = RewardConfig(4 output_dir="rm-out",5 max_length=1024,6 num_train_epochs=1, # one. Reward models overfit fast7 learning_rate=1e-5,8 per_device_train_batch_size=8,9 eval_strategy="steps", eval_steps=100,10 bf16=True,11)1213trainer = RewardTrainer(14 model=model, # AutoModelForSequenceClassification,15 args=cfg, # num_labels=116 processing_class=tok,17 train_dataset=ds, # columns: chosen, rejected18 eval_dataset=ds_eval,19)20trainer.train()The trainer implements the Bradley-Terry pairwise objective — push the chosen score above the rejected one, ease off once the gap is comfortable — and reports eval_accuracy, the fraction of held-out pairs ranked correctly. Two numbers it does not report and you must compute yourself: the correlation between reward and response length, and the accuracy of a trivial "always pick the longer response" baseline. Without those, an accuracy of 0.71 could mean a good reward model or a length detector.
Stage three, option A: DPO
1from trl import DPOTrainer, DPOConfig23cfg = DPOConfig(4 output_dir="dpo-out",5 beta=0.1, # KL strength; 0.05-0.56 learning_rate=5e-7, # 10-100x lower than SFT7 max_length=1024, # prompt + response, in tokens8 loss_type="sigmoid", # "ipo", "robust", "hinge", ...9 num_train_epochs=1,10 bf16=True,11)1213trainer = DPOTrainer(14 model=policy, ref_model=None, # None + LoRA: adapters off == ref15 args=cfg, processing_class=tok,16 train_dataset=ds, # columns: prompt, chosen, rejected17 peft_config=lora_cfg,18)19trainer.train()Watch rewards/accuracies and rewards/margins in the logs. A run that is working starts with a loss of exactly 0.6931 — at step zero the policy and reference are identical, so every margin is zero and the loss is log2. If your first logged loss is anything else, the reference is not a frozen copy of the starting policy and the run is meaningless.
Stage three, option B: online RL with a reward function
TRL 1.x has no PPO trainer. Its online trainers are critic-free: GRPOTrainer and RLOOTrainer sample several completions per prompt, score them with your reward functions, and use the group as the baseline. Here GRPO trains a small model on GSM8K maths problems against a correctness checker, the RL-with-verifiable-rewards setup:
1import re2from datasets import load_dataset3from trl import GRPOTrainer, GRPOConfig45# Every row needs a "prompt". Extra columns (here "answer") are passed6# to each reward function as keyword arguments.7ds = load_dataset("openai/gsm8k", "main", split="train")8ds = ds.map(lambda ex: {9 "prompt": [{"role": "user", "content": ex["question"]}],10 "answer": ex["answer"].split("####")[-1].strip().replace(",", ""),11})1213def correct(completions, answer, **kwargs):14 """1.0 if the last number in the reply equals the reference answer."""15 scores = []16 for msg, ref in zip(completions, answer):17 nums = re.findall(r"-?\d+(?:\.\d+)?", msg[0]["content"].replace(",", ""))18 scores.append(1.0 if nums and nums[-1] == ref else 0.0)19 return scores2021cfg = GRPOConfig(22 output_dir="grpo-out",23 num_generations=8, # the group: samples per prompt24 max_completion_length=512,25 learning_rate=1e-6,26 beta=0.0, # the default: no KL term, no reference model27 bf16=True,28)29trainer = GRPOTrainer(model="Qwen/Qwen2.5-0.5B-Instruct",30 reward_funcs=correct, args=cfg, train_dataset=ds)31trainer.train()Note that the reward is an arbitrary Python function you write. It does not have to come from a reward model — it can be a unit-test pass rate, a regex match, a verifier's verdict, or a list of several functions combined. This is the flexibility online RL buys and preference-only methods cannot offer. You can also pass a reward-model checkpoint instead of a function.
Two defaults in current TRL are worth knowing. beta is 0.0, so there is no KL penalty and no reference model is loaded; set it above zero (for example 0.04) when you want the leash back, as you usually do with a learned reward model. And the default loss_type is "dapo", which averages the loss over all tokens in the batch rather than per sequence, to remove the length bias of the original GRPO loss. Watch frac_reward_zero_std in the logs: it is the share of prompts where every sample got the same reward, which teach the model nothing. If it is near 1.0, your prompts are too easy or too hard for the model.
Axolotl: the configuration-first approach
Axolotl inverts the interface. Instead of writing Python, you write one YAML file that fully describes the run, and the CLI executes it. The consequence that matters is reproducibility: the config is the experiment, it goes in version control, and a colleague can rerun it exactly.
pip install --no-build-isolation "axolotl[deepspeed]"axolotl fetch examples # optional: copies the example configs# qlora-sft.ymlbase_model: meta-llama/Llama-3.2-3Bmodel_type: AutoModelForCausalLMload_in_4bit: true # QLoRA: 4-bit base weightsadapter: qloralora_r: 32lora_alpha: 64lora_dropout: 0.05lora_target_modules: - q_proj - k_proj - v_proj - o_projdatasets: - path: tatsu-lab/alpaca type: alpaca # built-in prompt formats; also "sharegpt", # "chat_template", or a custom mappingdataset_prepared_path: ./preparedval_set_size: 0.05sequence_len: 2048sample_packing: truepad_to_sequence_len: truemicro_batch_size: 2gradient_accumulation_steps: 8 # effective batch = 16 per devicenum_epochs: 3learning_rate: 0.0002lr_scheduler: cosinewarmup_ratio: 0.03optimizer: paged_adamw_8bitbf16: autoflash_attention: truegradient_checkpointing: trueoutput_dir: ./out-qloralogging_steps: 10save_strategy: epoch1# Tokenise and cache the dataset first - fails fast on format errors2axolotl preprocess qlora-sft.yml34# Then train (uses all visible GPUs; --launcher picks accelerate or torchrun)5axolotl train qlora-sft.yml67# Merge the adapter into the base weights for deployment8axolotl merge-lora qlora-sft.yml --lora-model-dir=./out-qloraPreference training changes a handful of lines rather than the whole script:
# dpo.yml - the diff from the SFT configrl: dporl_beta: 0.1learning_rate: 0.000005 # much lower than SFTdatasets: - path: Intel/orca_dpo_pairs type: chatml.intel # maps the dataset's columns to # prompt / chosen / rejected# Start from the SFT adapter you just trainedbase_model: meta-llama/Llama-3.2-3Badapter: loralora_model_dir: ./out-qloraThe type: field on a dataset is where most Axolotl time is lost. It selects a built-in mapping from that dataset's column names and conversation structure into the prompt format the model expects. Choose the wrong one and training runs perfectly while learning a format the model will never see at inference. Always run preprocess first (add --debug to print decoded examples) and read a sample before launching a long job.
Ask one question: is anything about this run non-standard? If the answer is no, take the config-driven path and spend the saved days on your data instead.
Choosing between them
| Situation | Use | Reason |
|---|---|---|
| Standard SFT or DPO on a supported model, want it running today | Axolotl | One config file, no Python; sensible defaults for memory and packing |
| Custom reward function — unit tests, a verifier, a weighted mixture | TRL | The reward is a Python callable you control |
| Custom loss, or a research variant not yet in any library | TRL | Trainers are subclassable; the loss is a method you can override |
| PPO at large scale where rollout generation dominates | OpenRLHF or veRL | Ray plus vLLM for rollouts; generation is the bottleneck and it is optimised there |
| Single consumer GPU, tight memory | Unsloth, or Axolotl with QLoRA | Fused kernels and 4-bit base weights |
| Many model families, prefer a UI | LLaMA-Factory | Broadest architecture coverage; web interface for configuration |
| Reproducibility across a team | Axolotl | The config is the complete specification of the run |
These are not exclusive. A very common arrangement is Axolotl for the SFT stage because it is fast to configure, TRL for the reward model and PPO stages because they need custom Python, and a shared evaluation harness across both.
Sizing the run before you launch it
Most wasted GPU hours come from launching a job that could not have fitted, discovering it forty minutes in, and repeating. The arithmetic is simple enough to do in advance.
Full fine-tuning in bf16 with Adam costs roughly 16 bytes per parameter: 2 for the weights, 2 for the gradient, and 12 for the optimiser's two moments plus an fp32 master copy. A 7B model is therefore about 112 GB before a single activation. LoRA changes the picture completely, because gradients and optimiser states exist only for the adapter — a few tens of millions of parameters instead of seven billion.
| Setup | 7B model, approximate | Fits on | Trade-off |
|---|---|---|---|
| Full fine-tune, bf16 + Adam | ~112 GB + activations | 2–4 × 80 GB with sharding | Best quality; expensive |
| LoRA, bf16 base | ~14 GB base + ~1 GB adapter states | 1 × 40 GB | Close to full quality on most alignment tasks |
| QLoRA, 4-bit base | ~4 GB base + adapter states | 1 × 24 GB | Slower per step; small quality cost |
| DPO with LoRA | Same as LoRA — the reference is the base with adapters off | 1 × 40 GB | The reference model is genuinely free |
| PPO with LoRA | LoRA policy + value head + a separate reward model | 1 × 80 GB, tight | Generation in the loop dominates wall-clock time |
The DPO row is the one worth internalising. With adapters, the frozen reference is the same base weights with the adapters switched off, so the second model in the DPO loss costs nothing but an extra forward pass. That single fact is why DPO with LoRA fits comfortably where a naive reading of "you need two models" suggests it should not.
Methods newer than the classical pipeline
TRL has kept adding trainers as the field has moved, and two families are worth knowing about because they change what the pipeline looks like.
- Group-relative methods (
GRPOTrainer,RLOOTrainer) drop the value model entirely. Instead of a learned critic predicting a baseline, they sample a group of responses to the same prompt and use the group's mean reward as the baseline (GRPO also divides by the group's standard deviation). That removes one trainable model from memory and one source of instability — value divergence — at the cost of needing several samples per prompt. Where a programmatic reward exists, such as a unit-test pass rate, this has become the common choice, and it is how most open reasoning models since DeepSeek-R1 have been trained. - Online preference methods (
OnlineDPOTrainer,NashMDTrainer, now intrl.experimental) sample from the current policy during training and label those samples with a judge, rather than consuming a fixed dataset. They recover the on-policy property that plain DPO lacks, without a value model or a clipped objective.
Both are configured the same way as everything above — a config object, a dataset, a trainer — which is the practical benefit of a library that has kept one interface across five years of method churn.
Chaining the stages
1#!/usr/bin/env bash2set -euo pipefail # stop on the first failure - a pipeline that3 # continues past a failed stage wastes hours45# 1. SFT with Axolotl6axolotl preprocess configs/sft.yml7axolotl train configs/sft.yml8axolotl merge-lora configs/sft.yml --lora-model-dir=./out-sft910# 2. Reward model with TRL (needs a custom eval, so Python)11python scripts/train_reward_model.py \12 --base ./out-sft-merged --data data/prefs --out ./out-rm1314# 2b. GATE: refuse to continue on a bad reward model15python scripts/audit_rm.py --rm ./out-rm --data data/prefs_test \16 --min-accuracy 0.65 --max-length-corr 0.251718# 3. Policy optimisation19python scripts/train_dpo.py \20 --policy ./out-sft-merged --data data/prefs --beta 0.1 --out ./out-dpo2122# 4. Evaluate: capability, safety, and a head-to-head win rate23lm_eval --model hf --model_args pretrained=./out-dpo \24 --tasks mmlu,truthfulqa_mc2,hellaswag --batch_size 8 \25 --output_path results/dpo.json26python scripts/win_rate.py --a ./out-sft-merged --b ./out-dpo \27 --prompts data/heldout_prompts.jsonlStep 2b is the part teams leave out and then regret. A reward model that scores 0.58 accuracy with a 0.4 length correlation will happily drive a policy for eight hours and produce something worse than what you started with. A gate that costs two minutes of evaluation prevents that, and it belongs in the script rather than in someone's memory.
Evaluating with lm-evaluation-harness
1pip install lm-eval23lm_eval --model hf \4 --model_args pretrained=./out-dpo,dtype=bfloat16 \5 --tasks mmlu,gsm8k,truthfulqa_mc2,toxigen \6 --num_fewshot 5 --batch_size auto \7 --output_path results/after.json89# Always run the SAME command on the model you STARTED from.10# A benchmark number without a baseline is not information.11lm_eval --model hf \12 --model_args pretrained=./out-sft-merged,dtype=bfloat16 \13 --tasks mmlu,gsm8k,truthfulqa_mc2,toxigen \14 --num_fewshot 5 --batch_size auto \15 --output_path results/before.jsonThe comparison is the point. Alignment training typically costs one to four points of MMLU — the alignment tax — and the only way to know whether you paid two points or fifteen is to have measured the same tasks, with the same few-shot count and the same harness version, before and after.
Benchmarks also will not tell you whether the model is more helpful. For that you need a head-to-head win rate: generate responses from both models on held-out prompts, present them blind and in randomised order, and count. A judge model can stand in for human raters, provided you average over both presentation orders to cancel position bias.
A benchmark number with no before-measurement is not a result. Run the identical evaluation on the model you started from, or you cannot tell an improvement from a regression.
Failures that are specific to the tooling
| Symptom | Cause | Fix |
|---|---|---|
TypeError: unexpected keyword argument on a trainer | Library version differs from the tutorial you copied | Check the installed version's own documentation; pin versions in the repo |
| Loss falls, generations are nonsense or in the wrong format | Chat template mismatch between training and inference | Decode a prepared training sample and compare it byte-for-byte with what you send at inference |
| DPO first loss is not 0.6931 | Reference model is not a frozen copy of the starting policy | Verify checkpoints match and the reference is in eval() under no_grad() |
| Model invents a follow-up user turn after answering | sample_packing without proper cross-example attention masking | Disable packing, or verify the masking implementation for your model type |
| PPO generations are truncated or start mid-word | Right padding used for generation; must be left | Set tokenizer.padding_side = "left" for the generation path |
| OOM only after several hundred steps | Rollout buffer or logged tensors accumulating on device | Detach and move logged tensors to CPU; cap the buffer explicitly |
| Multi-GPU run gives different results than single-GPU | Effective batch size differs — it is micro-batch × accumulation × world size | Compute the effective batch and adjust accumulation when changing GPU count |
| Merged LoRA model behaves worse than the adapter did | Merged in a different dtype than trained, or merged onto a different base revision | Merge in the training dtype; pin the base model revision hash |
What this means when you set up a project
Pin every version on day one and record them in the repository. This ecosystem moves fast enough that a working configuration is a genuine artefact. A requirements.txt with exact versions, committed alongside your configs, is the difference between a reproducible result and one you cannot recreate in four months.
Decode a prepared sample before every long run. One minute, and it catches the entire family of format bugs — wrong chat template, wrong dataset type mapping, missing end-of-turn token, prompt masked incorrectly — that otherwise present as a training run that looks perfect and produces a worse model.
Debug the whole pipeline at toy scale first. A 125M model, 500 examples, 50 steps, one GPU. Every structural bug appears here in minutes: tokeniser mismatch, unfrozen reference, reward shape errors, packing artefacts. Scaling a working pipeline is routine; debugging a broken one at 7B across eight GPUs is not.
Put a quality gate between every stage. Minimum reward-model accuracy, maximum length correlation, minimum DPO reward accuracy, maximum MMLU regression. Encode them as assertions in the pipeline script that halt the run. Every one of these has a threshold you already know; the only question is whether the machine checks it or a person remembers to.
Choose the framework by what is non-standard about your run. If nothing is — standard model, standard method, standard data format — take the config-driven path and spend your saved time on the data. If your reward is computed by code you wrote, or your loss is not in any library, you need the programmatic one. That single question decides it more reliably than any feature comparison.