Reinforcement Learning from Human Feedback (RLHF)

Open-Source RLHF Frameworks — TRL, Axolotl, TRLX


An engineer decides to write PPO from scratch. It is not unreasonable — the clipped objective is about twenty lines, GAE is another fifteen, and the paper is clear. Three days later it runs correctly on a 125M model on one GPU. The reward goes up. Everything works.

Then it has to run on a 7B model, and the next three weeks look like this. It runs out of memory, so ZeRO stage 3 shards the optimiser states — and now model.generate() silently produces garbage, because the parameters are sharded across ranks and generation needs them gathered. Gradient checkpointing halves activation memory but conflicts with the value head's forward hook. Mixed precision makes the KL estimate unstable at bf16. Multi-GPU rollouts need the batch resharded between the generation phase and the optimisation phase. The tokeniser pads right for training and needs to pad left for generation, and getting that wrong shifts every log-probability by one position without raising an error.

None of that is PPO. All of it is distributed-systems plumbing, and every one of those problems has already been solved, argued about, and fixed in a library.

The value of an RLHF framework is almost never the loss function. It is that someone has already reconciled sharded parameters with autoregressive generation, and you did not have to.

What each framework actually gives youPython APIYAML configPython APIYes, SFTTrainerYes, first classNot the focusYes, RewardTrainerVia TRL under itYesYes, DPOTrainerYes, oneconfig keyNoYes, PPOTrainerDelegates to TRLYes, distributedActively developedActively developedLargely dormantTRLAxolotlTRLXInterfaceSFTReward modelDPOPPOUpkeep
Axolotl is a configuration layer over TRL rather than a rival to it, so the real choice is whether you want the stages as code or as YAML.

The landscape

ToolInterfaceCoversScale ceilingStatus
TRL (Hugging Face)Python APISFT, reward models, DPO and variants, KTO, GRPO, RLOO; more in trl.experimental. No PPO in 1.xSingle node to modest multi-nodeActively developed; the default for most work
AxolotlYAML config + CLISFT, LoRA/QLoRA, DPO, ORPO, KTO, GRPOMulti-GPU, multi-node via DeepSpeed/FSDPActive; the fastest path from data to a trained adapter
TRLX (CarperAI)Python APIPPO, ILQLWas the multi-node option of its eraEffectively unmaintained — historically important, not a choice for new work
LLaMA-FactoryYAML + web UISFT, PPO, DPO, KTO, many model familiesMulti-GPUActive; broadest model support
OpenRLHFPython + shell scriptsPPO, REINFORCE++, GRPO, RLOO, plus SFT, reward models and DPOLarge multi-node; Ray-based, vLLM for rolloutsActive; built for the scale where rollout generation dominates
veRL (ByteDance)Python + Hydra configPPO, GRPO, DAPO, RLOO, REINFORCE++ and moreLarge multi-node; FSDP or Megatron for training, vLLM or SGLang for rolloutsActive; widely used for RL on reasoning models
UnslothPython APISFT, DPO, GRPO with custom fused kernelsSingle GPUActive; large speed and memory wins on one card

Two of these are worth knowing properly, because between them they cover most real work: TRL when you need programmatic control, Axolotl when you need a reproducible pipeline fast. For PPO with a critic, or RL at a scale where generation dominates, use OpenRLHF or veRL; section 2 shows an OpenRLHF launch.

A note on TRLX

TRLX mattered. When RLHF first became reproducible outside large labs, it was the library that made distributed PPO on multi-billion-parameter models practical, and it introduced the pattern of a reward function passed in as a plain Python callable:

Python
# The TRLX idea, which everything since has borrowed:# training is driven by an arbitrary function from text to float.def reward_fn(samples, prompts=None, outputs=None):    return [float(score(rm, rm_tok, s)) for s in samples]# trlx.train(reward_fn=reward_fn, prompts=prompts, config=cfg)

That callable interface is the right abstraction and it survives in every modern library. The project itself, however, has not kept pace with current model architectures and training stacks. Treat it as background: if you inherit a TRLX codebase, understand it; do not start one.

TRL in practice

Bash
pip install "trl>=1.0,<2" "transformers>=5" peft accelerate \            datasets bitsandbytes# Pin these. The code below was checked against trl 1.13. TRL's# trainers have changed across versions more than once (PPOTrainer# was removed; max_prompt_length left DPOConfig), and code from an# older tutorial frequently fails on a renamed argument.pip freeze | grep -E "^(trl|transformers|peft|accelerate)=="

Stage one: supervised fine-tuning

Python
from trl import SFTTrainer, SFTConfigfrom peft import LoraConfigcfg = SFTConfig(    output_dir="sft-out",    max_length=2048,    packing=True,                 # concatenate short examples to fill                                  # the context - large throughput win,                                  # but see the warning below    num_train_epochs=2,    per_device_train_batch_size=4,    gradient_accumulation_steps=4,    learning_rate=2e-4,           # high because LoRA; 2e-5 for full FT    bf16=True,)trainer = SFTTrainer(    model="meta-llama/Llama-3.2-1B",    args=cfg,    train_dataset=ds,             # column: "text" or "messages"    peft_config=LoraConfig(r=16, lora_alpha=32, lora_dropout=0.05,                           target_modules=["q_proj", "k_proj",                                           "v_proj", "o_proj"]),)trainer.train()

packing=True deserves the warning. It concatenates multiple short examples into one sequence to avoid wasting context on padding, which can double throughput on datasets of short turns. It also means examples bleed across boundaries unless the implementation masks attention between them. For instruction data with short responses this occasionally teaches the model to continue past its own end-of-turn token into the next example. If your fine-tuned model starts inventing a follow-up user turn after answering, packing is the first thing to switch off.

Stage two: the reward model

Python
from trl import RewardTrainer, RewardConfigcfg = RewardConfig(    output_dir="rm-out",    max_length=1024,    num_train_epochs=1,           # one. Reward models overfit fast    learning_rate=1e-5,    per_device_train_batch_size=8,    eval_strategy="steps", eval_steps=100,    bf16=True,)trainer = RewardTrainer(    model=model,                  # AutoModelForSequenceClassification,    args=cfg,                     # num_labels=1    processing_class=tok,    train_dataset=ds,             # columns: chosen, rejected    eval_dataset=ds_eval,)trainer.train()

The trainer implements the Bradley-Terry pairwise objective — push the chosen score above the rejected one, ease off once the gap is comfortable — and reports eval_accuracy, the fraction of held-out pairs ranked correctly. Two numbers it does not report and you must compute yourself: the correlation between reward and response length, and the accuracy of a trivial "always pick the longer response" baseline. Without those, an accuracy of 0.71 could mean a good reward model or a length detector.

Stage three, option A: DPO

Python
from trl import DPOTrainer, DPOConfigcfg = DPOConfig(    output_dir="dpo-out",    beta=0.1,                     # KL strength; 0.05-0.5    learning_rate=5e-7,           # 10-100x lower than SFT    max_length=1024,              # prompt + response, in tokens    loss_type="sigmoid",          # "ipo", "robust", "hinge", ...    num_train_epochs=1,    bf16=True,)trainer = DPOTrainer(    model=policy, ref_model=None, # None + LoRA: adapters off == ref    args=cfg, processing_class=tok,    train_dataset=ds,             # columns: prompt, chosen, rejected    peft_config=lora_cfg,)trainer.train()

Watch rewards/accuracies and rewards/margins in the logs. A run that is working starts with a loss of exactly 0.6931 — at step zero the policy and reference are identical, so every margin is zero and the loss is log⁡2\log 2. If your first logged loss is anything else, the reference is not a frozen copy of the starting policy and the run is meaningless.

Stage three, option B: online RL with a reward function

TRL 1.x has no PPO trainer. Its online trainers are critic-free: GRPOTrainer and RLOOTrainer sample several completions per prompt, score them with your reward functions, and use the group as the baseline. Here GRPO trains a small model on GSM8K maths problems against a correctness checker, the RL-with-verifiable-rewards setup:

Python
import refrom datasets import load_datasetfrom trl import GRPOTrainer, GRPOConfig# Every row needs a "prompt". Extra columns (here "answer") are passed# to each reward function as keyword arguments.ds = load_dataset("openai/gsm8k", "main", split="train")ds = ds.map(lambda ex: {    "prompt": [{"role": "user", "content": ex["question"]}],    "answer": ex["answer"].split("####")[-1].strip().replace(",", ""),})def correct(completions, answer, **kwargs):    """1.0 if the last number in the reply equals the reference answer."""    scores = []    for msg, ref in zip(completions, answer):        nums = re.findall(r"-?\d+(?:\.\d+)?", msg[0]["content"].replace(",", ""))        scores.append(1.0 if nums and nums[-1] == ref else 0.0)    return scorescfg = GRPOConfig(    output_dir="grpo-out",    num_generations=8,            # the group: samples per prompt    max_completion_length=512,    learning_rate=1e-6,    beta=0.0,                     # the default: no KL term, no reference model    bf16=True,)trainer = GRPOTrainer(model="Qwen/Qwen2.5-0.5B-Instruct",                      reward_funcs=correct, args=cfg, train_dataset=ds)trainer.train()

Note that the reward is an arbitrary Python function you write. It does not have to come from a reward model — it can be a unit-test pass rate, a regex match, a verifier's verdict, or a list of several functions combined. This is the flexibility online RL buys and preference-only methods cannot offer. You can also pass a reward-model checkpoint instead of a function.

Two defaults in current TRL are worth knowing. beta is 0.0, so there is no KL penalty and no reference model is loaded; set it above zero (for example 0.04) when you want the leash back, as you usually do with a learned reward model. And the default loss_type is "dapo", which averages the loss over all tokens in the batch rather than per sequence, to remove the length bias of the original GRPO loss. Watch frac_reward_zero_std in the logs: it is the share of prompts where every sample got the same reward, which teach the model nothing. If it is near 1.0, your prompts are too easy or too hard for the model.

Axolotl: the configuration-first approach

Axolotl inverts the interface. Instead of writing Python, you write one YAML file that fully describes the run, and the CLI executes it. The consequence that matters is reproducibility: the config is the experiment, it goes in version control, and a colleague can rerun it exactly.

Bash
pip install --no-build-isolation "axolotl[deepspeed]"axolotl fetch examples        # optional: copies the example configs
Text
# qlora-sft.ymlbase_model: meta-llama/Llama-3.2-3Bmodel_type: AutoModelForCausalLMload_in_4bit: true          # QLoRA: 4-bit base weightsadapter: qloralora_r: 32lora_alpha: 64lora_dropout: 0.05lora_target_modules:  - q_proj  - k_proj  - v_proj  - o_projdatasets:  - path: tatsu-lab/alpaca    type: alpaca            # built-in prompt formats; also "sharegpt",                            # "chat_template", or a custom mappingdataset_prepared_path: ./preparedval_set_size: 0.05sequence_len: 2048sample_packing: truepad_to_sequence_len: truemicro_batch_size: 2gradient_accumulation_steps: 8      # effective batch = 16 per devicenum_epochs: 3learning_rate: 0.0002lr_scheduler: cosinewarmup_ratio: 0.03optimizer: paged_adamw_8bitbf16: autoflash_attention: truegradient_checkpointing: trueoutput_dir: ./out-qloralogging_steps: 10save_strategy: epoch
Bash
# Tokenise and cache the dataset first - fails fast on format errorsaxolotl preprocess qlora-sft.yml# Then train (uses all visible GPUs; --launcher picks accelerate or torchrun)axolotl train qlora-sft.yml# Merge the adapter into the base weights for deploymentaxolotl merge-lora qlora-sft.yml --lora-model-dir=./out-qlora

Preference training changes a handful of lines rather than the whole script:

Text
# dpo.yml - the diff from the SFT configrl: dporl_beta: 0.1learning_rate: 0.000005      # much lower than SFTdatasets:  - path: Intel/orca_dpo_pairs    type: chatml.intel        # maps the dataset's columns to                              # prompt / chosen / rejected# Start from the SFT adapter you just trainedbase_model: meta-llama/Llama-3.2-3Badapter: loralora_model_dir: ./out-qlora

The type: field on a dataset is where most Axolotl time is lost. It selects a built-in mapping from that dataset's column names and conversation structure into the prompt format the model expects. Choose the wrong one and training runs perfectly while learning a format the model will never see at inference. Always run preprocess first (add --debug to print decoded examples) and read a sample before launching a long job.

Ask one question: is anything about this run non-standard? If the answer is no, take the config-driven path and spend the saved days on your data instead.

Choosing between them

SituationUseReason
Standard SFT or DPO on a supported model, want it running todayAxolotlOne config file, no Python; sensible defaults for memory and packing
Custom reward function — unit tests, a verifier, a weighted mixtureTRLThe reward is a Python callable you control
Custom loss, or a research variant not yet in any libraryTRLTrainers are subclassable; the loss is a method you can override
PPO at large scale where rollout generation dominatesOpenRLHF or veRLRay plus vLLM for rollouts; generation is the bottleneck and it is optimised there
Single consumer GPU, tight memoryUnsloth, or Axolotl with QLoRAFused kernels and 4-bit base weights
Many model families, prefer a UILLaMA-FactoryBroadest architecture coverage; web interface for configuration
Reproducibility across a teamAxolotlThe config is the complete specification of the run

These are not exclusive. A very common arrangement is Axolotl for the SFT stage because it is fast to configure, TRL for the reward model and PPO stages because they need custom Python, and a shared evaluation harness across both.

Sizing the run before you launch it

Most wasted GPU hours come from launching a job that could not have fitted, discovering it forty minutes in, and repeating. The arithmetic is simple enough to do in advance.

Full fine-tuning in bf16 with Adam costs roughly 16 bytes per parameter: 2 for the weights, 2 for the gradient, and 12 for the optimiser's two moments plus an fp32 master copy. A 7B model is therefore about 112 GB before a single activation. LoRA changes the picture completely, because gradients and optimiser states exist only for the adapter — a few tens of millions of parameters instead of seven billion.

Setup7B model, approximateFits onTrade-off
Full fine-tune, bf16 + Adam~112 GB + activations2–4 × 80 GB with shardingBest quality; expensive
LoRA, bf16 base~14 GB base + ~1 GB adapter states1 × 40 GBClose to full quality on most alignment tasks
QLoRA, 4-bit base~4 GB base + adapter states1 × 24 GBSlower per step; small quality cost
DPO with LoRASame as LoRA — the reference is the base with adapters off1 × 40 GBThe reference model is genuinely free
PPO with LoRALoRA policy + value head + a separate reward model1 × 80 GB, tightGeneration in the loop dominates wall-clock time

The DPO row is the one worth internalising. With adapters, the frozen reference is the same base weights with the adapters switched off, so the second model in the DPO loss costs nothing but an extra forward pass. That single fact is why DPO with LoRA fits comfortably where a naive reading of "you need two models" suggests it should not.

Methods newer than the classical pipeline

TRL has kept adding trainers as the field has moved, and two families are worth knowing about because they change what the pipeline looks like.

  • Group-relative methods (GRPOTrainer, RLOOTrainer) drop the value model entirely. Instead of a learned critic predicting a baseline, they sample a group of responses to the same prompt and use the group's mean reward as the baseline (GRPO also divides by the group's standard deviation). That removes one trainable model from memory and one source of instability — value divergence — at the cost of needing several samples per prompt. Where a programmatic reward exists, such as a unit-test pass rate, this has become the common choice, and it is how most open reasoning models since DeepSeek-R1 have been trained.
  • Online preference methods (OnlineDPOTrainer, NashMDTrainer, now in trl.experimental) sample from the current policy during training and label those samples with a judge, rather than consuming a fixed dataset. They recover the on-policy property that plain DPO lacks, without a value model or a clipped objective.

Both are configured the same way as everything above — a config object, a dataset, a trainer — which is the practical benefit of a library that has kept one interface across five years of method churn.

Chaining the stages

Bash
#!/usr/bin/env bashset -euo pipefail        # stop on the first failure - a pipeline that                         # continues past a failed stage wastes hours# 1. SFT with Axolotlaxolotl preprocess configs/sft.ymlaxolotl train configs/sft.ymlaxolotl merge-lora configs/sft.yml --lora-model-dir=./out-sft# 2. Reward model with TRL (needs a custom eval, so Python)python scripts/train_reward_model.py \    --base ./out-sft-merged --data data/prefs --out ./out-rm# 2b. GATE: refuse to continue on a bad reward modelpython scripts/audit_rm.py --rm ./out-rm --data data/prefs_test \    --min-accuracy 0.65 --max-length-corr 0.25# 3. Policy optimisationpython scripts/train_dpo.py \    --policy ./out-sft-merged --data data/prefs --beta 0.1 --out ./out-dpo# 4. Evaluate: capability, safety, and a head-to-head win ratelm_eval --model hf --model_args pretrained=./out-dpo \        --tasks mmlu,truthfulqa_mc2,hellaswag --batch_size 8 \        --output_path results/dpo.jsonpython scripts/win_rate.py --a ./out-sft-merged --b ./out-dpo \    --prompts data/heldout_prompts.jsonl

Step 2b is the part teams leave out and then regret. A reward model that scores 0.58 accuracy with a 0.4 length correlation will happily drive a policy for eight hours and produce something worse than what you started with. A gate that costs two minutes of evaluation prevents that, and it belongs in the script rather than in someone's memory.

Evaluating with lm-evaluation-harness

Bash
pip install lm-evallm_eval --model hf \        --model_args pretrained=./out-dpo,dtype=bfloat16 \        --tasks mmlu,gsm8k,truthfulqa_mc2,toxigen \        --num_fewshot 5 --batch_size auto \        --output_path results/after.json# Always run the SAME command on the model you STARTED from.# A benchmark number without a baseline is not information.lm_eval --model hf \        --model_args pretrained=./out-sft-merged,dtype=bfloat16 \        --tasks mmlu,gsm8k,truthfulqa_mc2,toxigen \        --num_fewshot 5 --batch_size auto \        --output_path results/before.json

The comparison is the point. Alignment training typically costs one to four points of MMLU — the alignment tax — and the only way to know whether you paid two points or fifteen is to have measured the same tasks, with the same few-shot count and the same harness version, before and after.

Benchmarks also will not tell you whether the model is more helpful. For that you need a head-to-head win rate: generate responses from both models on held-out prompts, present them blind and in randomised order, and count. A judge model can stand in for human raters, provided you average over both presentation orders to cancel position bias.

A benchmark number with no before-measurement is not a result. Run the identical evaluation on the model you started from, or you cannot tell an improvement from a regression.

Failures that are specific to the tooling

SymptomCauseFix
TypeError: unexpected keyword argument on a trainerLibrary version differs from the tutorial you copiedCheck the installed version's own documentation; pin versions in the repo
Loss falls, generations are nonsense or in the wrong formatChat template mismatch between training and inferenceDecode a prepared training sample and compare it byte-for-byte with what you send at inference
DPO first loss is not 0.6931Reference model is not a frozen copy of the starting policyVerify checkpoints match and the reference is in eval() under no_grad()
Model invents a follow-up user turn after answeringsample_packing without proper cross-example attention maskingDisable packing, or verify the masking implementation for your model type
PPO generations are truncated or start mid-wordRight padding used for generation; must be leftSet tokenizer.padding_side = "left" for the generation path
OOM only after several hundred stepsRollout buffer or logged tensors accumulating on deviceDetach and move logged tensors to CPU; cap the buffer explicitly
Multi-GPU run gives different results than single-GPUEffective batch size differs — it is micro-batch × accumulation × world sizeCompute the effective batch and adjust accumulation when changing GPU count
Merged LoRA model behaves worse than the adapter didMerged in a different dtype than trained, or merged onto a different base revisionMerge in the training dtype; pin the base model revision hash

What this means when you set up a project

Pin every version on day one and record them in the repository. This ecosystem moves fast enough that a working configuration is a genuine artefact. A requirements.txt with exact versions, committed alongside your configs, is the difference between a reproducible result and one you cannot recreate in four months.

Decode a prepared sample before every long run. One minute, and it catches the entire family of format bugs — wrong chat template, wrong dataset type mapping, missing end-of-turn token, prompt masked incorrectly — that otherwise present as a training run that looks perfect and produces a worse model.

Debug the whole pipeline at toy scale first. A 125M model, 500 examples, 50 steps, one GPU. Every structural bug appears here in minutes: tokeniser mismatch, unfrozen reference, reward shape errors, packing artefacts. Scaling a working pipeline is routine; debugging a broken one at 7B across eight GPUs is not.

Put a quality gate between every stage. Minimum reward-model accuracy, maximum length correlation, minimum DPO reward accuracy, maximum MMLU regression. Encode them as assertions in the pipeline script that halt the run. Every one of these has a threshold you already know; the only question is whether the machine checks it or a person remembers to.

Choose the framework by what is non-standard about your run. If nothing is — standard model, standard method, standard data format — take the config-driven path and spend your saved time on the data. If your reward is computed by code you wrote, or your loss is not in any library, you need the programmatic one. That single question decides it more reliably than any feature comparison.