Course Content
Fine-Tuning LLMs
6 sections · 52 lessons
What steps do you follow to build a good fine-tuning dataset?
What you need to know
- Write the spec — what a correct output is, with tricky cases decided. Hand-write 20–50 gold examples. If experts disagree on them, stop and fix the spec.
- Collect real inputs — production logs, tickets, documents. Invented inputs drift from what users actually send.
- Produce outputs — experts for the gold and test sets; a strong model plus human review for bulk training data.
- Clean — remove exact and near-duplicates, truncated or empty answers, wrong language, format violations and personal data you should not train on.
- Balance coverage — every intent and class, rare but important cases, and "I can't answer that" or "please clarify" examples.
- Split by entity — train, validation and test by customer, contract or patient, not by row.
- Check the rendering — apply the chat template, count tokens, and read 20 random rendered examples.
Format
Use one JSON object per line (JSONL) in the messages format from the instruction-tuning lesson, with the same system prompt you will use in production.
Quick checks in code
1from datasets import load_dataset2from transformers import AutoTokenizer34tok = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-7B-Instruct")5ds = load_dataset("json", data_files="train.jsonl", split="train")67def n_tokens(row):8 text = tok.apply_chat_template(row["messages"], tokenize=False)9 return {"n_tokens": len(tok(text).input_ids)}1011ds = ds.map(n_tokens)12print(max(ds["n_tokens"]), sum(t > 2048 for t in ds["n_tokens"]))This finds examples longer than your max_length (here 2,048), which would be silently truncated — often cutting off the answer, the part you most wanted the model to learn.
Why quality beats quantity
The LIMA paper (2023) showed that about 1,000 carefully chosen examples could produce a strong chat model from a good base model. The reason: fine-tuning mostly teaches style and format, and every inconsistent example teaches a contradiction. Ten thousand examples in which a quarter disagree with the spec give the model a noisy target.
Why split by entity
If the same contract appears in both train and test (as different clauses), the test score measures memory of that contract, not skill. Near-duplicates across splits are the most common reason for a great offline score and a disappointing launch.
A real-life example
The Mumbai law firm builds its clause-classifier dataset:
- Spec. Two partners disagree on whether "the vendor shall reimburse all losses" is indemnity or limitation of liability. They write a rule: any promise to pay for the other party's losses is indemnity; any cap on amounts is limitation. The spec now has 14 such rules.
- Inputs. 9,000 clauses from 1,200 past contracts, with client names and amounts masked.
- Cleaning. Near-duplicate removal finds that 35% of clauses are the firm's own standard boilerplate, repeated across contracts. They keep one copy of each, leaving about 5,900 clauses.
- Coverage. Arbitration clauses with a foreign seat are rare but high-risk, so associates label 150 more.
- Split. By contract: all clauses from one contract go to the same split.
The first fine-tune on this set beats an earlier one trained on 20,000 unfiltered clauses.
Follow-up questions to expect
- "How do you check label quality at scale?" — Double-label a sample and measure agreement; use an LLM judge with the spec as a rubric to flag suspicious rows; review the rows with the highest training loss.
- "Can you use synthetic data?" — Yes, for bulk outputs or rare cases, with filtering and human spot checks. Keep the test set real and human-labelled (section 2).
- "How much data do you need?" — Start with a few hundred, train, and look at the learning curve. Add data where the error analysis says the model is weak, not everywhere.