Course Content
Fine-Tuning LLMs
6 sections · 52 lessons
How is a preference dataset for reward modeling or DPO structured (chosen vs rejected, metadata)?
What you need to know
The record
1{2 "prompt": [{"role": "user", "content": "Mera refund kab aayega? Order 48213"}],3 "chosen": [{"role": "assistant", "content": "Aapka refund 12 June ko approve hua hai..."}],4 "rejected": [{"role": "assistant", "content": "Refunds usually take some time."}],5 "meta": {"source": "agent_edit", "annotator": "agent_117", "margin": "strong",6 "criteria": ["specific", "language_match"], "category": "refund"}7}Rules that matter:
- Same prompt for both answers. Comparing answers to different prompts teaches nothing.
- Chat format that matches how you will serve.
DPOTrainerapplies the model's chat template to conversational records. - Multi-turn: the shared history goes in
prompt; only the last assistant turn differs between chosen and rejected. - Reward models use the same format. TRL's
RewardTraineralso takes chosen/rejected pairs.
Metadata worth keeping
| Field | Why |
|---|---|
source (human, AI judge, agent edit, sampling) | Slice results by where pairs came from |
annotator, agreement | Find and drop unreliable raters |
margin or rating gap | Filter weak pairs, or weight strong ones |
criteria (why chosen won) | Know what the model is actually learning |
is_tie | Keep the flag, but leave ties out of training |
Where pairs come from
- Human ranking of two or more model answers.
- Production signals: a human edit (edited version is chosen), a regenerate click, a thumbs-down.
- AI judges scoring model answers against a rubric (RLAIF, covered in Section 3).
- On-policy sampling: generate 2–4 answers from your current SFT model and rank them. Pairs drawn from your own model usually work better than pairs from a different model, because the model learns about its own mistakes, not about another model's style.
Hygiene checks before training
1from datasets import load_dataset23ds = load_dataset("json", data_files="pairs.jsonl", split="train")4longer = sum(len(r["chosen"][-1]["content"]) > len(r["rejected"][-1]["content"]) for r in ds)5print(f"chosen is longer in {longer / len(ds):.0%} of pairs")If the chosen answer is longer 80% of the time, the easiest lesson for the model is "be longer", not "be better". Also drop ties and near-identical pairs, deduplicate prompts between train and eval, and keep a few hundred held-out pairs to measure how often the trained model prefers the chosen answer.
A real-life example
A telecom company's Hindi customer-support model drafts replies, and human agents edit them before sending. Each edit becomes a pair: the agent's version is chosen, the model's draft is rejected. In two months they collect 9,000 pairs.
The metadata saves them twice. First, 2,100 pairs have an edit distance under 5 characters — agents fixing a typo or a name. Those pairs teach nothing about quality, so they are dropped. Second, slicing by annotator shows one agent who rewrote every reply into formal English, even for customers writing Hinglish. Those 600 pairs would have taught the model to stop matching the customer's language, so they are removed. The final 6,300 pairs go into DPO.
Follow-up questions to expect
- "How many pairs do you need?" — A few thousand clean pairs is a common start; a few hundred can already shift tone. Quality and agreement matter more than count.
- "Can chosen and rejected come from different models?" — Yes, but if one model is always chosen, the model may learn that model's style rather than quality. Pairs from your own model are safer.
- "What's the difference between reward-model data and DPO data?" — The format is the same. A reward model turns pairs into a scoring function used later in RL; DPO uses the pairs directly to update the model.