Course Content
Fine-Tuning LLMs with LoRA, QLoRA and PEFT
4 sections · 10 lessons
Dataset Design, Quality and Labelling Strategies
A team fine-tuned a 7B model on 50,000 support conversations and shipped it. Within two days, users noticed something odd: the model ended almost every reply with "Is there anything else I can help you with today?" — even when the customer had just written "no thanks, that's all, goodbye".
The cause was in the data. Their export had included the agent's closing pleasantry as part of the target output on 43,000 of the 50,000 examples. The model learned the single most consistent pattern in the dataset, and it learned it perfectly. Loss curves looked excellent throughout.
Then they cleaned the data down to 6,000 examples: closings stripped, near-duplicates removed, three ambiguous ticket categories merged. The 6,000-example model beat the 50,000-example model on every metric they measured. This is the rule that governs everything below.
A model cannot be better than its data. It can only be a faithful, high-fidelity reproduction of whatever patterns your dataset actually contains — including the ones you did not intend to put there.
The quality hierarchy
When people ask "how much data do I need?", they are usually asking the wrong question third. The order of importance is:
- Correctness. Are the outputs actually right? A dataset with 15% wrong labels teaches the model to be wrong 15% of the time, in exactly the ways your labellers were wrong.
- Consistency. Do two examples of the same situation get the same treatment? Inconsistency is worse than error, because it teaches the model that the task is random.
- Coverage. Does the data span the real input distribution, including the awkward edges?
- Format. Is every example structurally identical in the ways that should not vary?
- Quantity. Last. Always last.
Doubling a dataset from 5,000 to 10,000 examples typically buys a small improvement. Fixing a 12% label error rate typically buys a large one. Teams spend their effort in the opposite proportion.
Defining the task before collecting anything
Write the specification first, on paper, in unambiguous language. It has three parts.
The task statement
One sentence naming the input, the output, and the decision rule. "Given the text of a customer support ticket, output exactly one of 14 category labels, choosing the category that describes the customer's primary request." That word "primary" is doing enormous work — without it, tickets mentioning both a refund and a delivery delay get labelled inconsistently by different annotators, and the model learns noise.
The input and output contract
Be explicit about what is included and what is not:
- Input: ticket subject and body, concatenated, truncated to 2,000 characters. Customer name removed. Prior conversation history excluded.
- Output: a lowercase snake_case category string from a fixed list of 14, nothing else — no explanation, no punctuation, no leading whitespace.
- Edge cases: if a ticket is spam, output
spam. If it is empty or unintelligible, outputneeds_review. Never invent a category.
Every ambiguity you leave in this document becomes label noise later, at roughly one confused annotator per ambiguity.
The intended distribution
Decide deliberately whether your training data should mirror production frequencies or be rebalanced. If 70% of real tickets are order_status, a mirrored dataset teaches the model a strong prior towards that class — which is genuinely useful at inference, but means the rare classes get very few examples. A common compromise: cap any single class at about 30% of the data and ensure every class has at least 200 examples, then keep your evaluation set at true production proportions so your metrics reflect reality.
How much data, actually
There is no formula that is right, but there are usable starting points. For LoRA fine-tuning of a 7B-class model:
| Task type | Workable minimum | Comfortable | Diminishing returns past |
|---|---|---|---|
| Output format enforcement (JSON schema, fixed layout) | 300 | 1,000-2,000 | ~5,000 |
| Classification, per class | 100 | 500 | ~2,000 per class |
| Style and tone adaptation | 500 | 2,000-5,000 | ~20,000 |
| Domain instruction-following | 2,000 | 10,000-30,000 | ~100,000 |
| Complex reasoning with intermediate steps | 5,000 | 50,000+ | Rarely reached |
A rough sizing heuristic for classification: examples ≈ number_of_classes × 300 × difficulty_factor, with difficulty 1 for well-separated classes and 3 for classes that human annotators frequently confuse. Fourteen support categories with moderate confusion gives 14 × 300 × 2 = 8,400 examples. Treat that as an order-of-magnitude estimate, not a target.
The honest procedure is empirical: train on 25%, 50% and 100% of what you have and plot the evaluation metric. If the curve is still climbing steeply at 100%, collect more. If it flattened between 50% and 100%, more data will not help and you should be fixing quality instead.
Getting the data
Mining what you already have
Almost always the best source. Resolved support tickets, historical reports, merged pull requests, closed case files. Real inputs paired with human-approved outputs, at zero labelling cost. The work is in cleaning: stripping signatures, removing boilerplate, filtering out the cases the humans got wrong.
Synthetic generation
Use a strong model to produce examples. This works well for input diversity and badly for ground truth. The useful pattern is to generate varied inputs synthetically and have humans supply or verify the outputs:
1SEED_PROMPT = """You are generating realistic customer support tickets2for an e-commerce company. Produce a ticket about: {topic}34Vary the writing style. Some customers are terse, some are angry,5some ramble. Include typos in roughly one in four tickets.6Output only the ticket text, 30-120 words."""78topics = [9 "a parcel marked delivered but not received",10 "a refund that has not appeared after 10 days",11 "a discount code that will not apply at checkout",12 # ...13]1415# Generate inputs synthetically, but have a human assign the label.16# Never let the generator both write the question and grade it.Two rules keep synthetic data from poisoning a run. First, cap it: if more than about half your dataset is synthetic, the model starts learning the generator's quirks rather than your task. Second, tag every example with its provenance so you can measure whether synthetic examples help or hurt.
Augmentation
Cheap ways to expand a small dataset, in rough order of safety:
| Technique | What it does | Risk |
|---|---|---|
| Paraphrasing the input | Same meaning, different wording | Low — but verify the label still holds |
| Instruction rewording | Same task, different phrasing of the prompt | Low; improves robustness |
| Entity substitution | Swap names, products, dates | Low, if the label does not depend on them |
| Typo and casing injection | Matches messy real input | Low; do it on ~15% of examples |
| Back-translation | Translate out and back for variety | Medium — can break domain terminology |
| Random word deletion or shuffling | Adds noise | High for generation tasks; usually harmful |
Critical rule: augment after splitting. If you paraphrase first and then split, a paraphrase of a training example lands in your test set and your metrics become fiction.
Labelling: getting agreement you can trust
One annotator versus several
A single annotator is cheap and perfectly self-consistent — and encodes one person's idiosyncratic reading of every edge case, with no way to detect it. Multiple annotators cost more but let you measure whether the task is even well defined.
The measurement is inter-annotator agreement, and raw agreement percentage is misleading because two annotators agree by chance. Cohen's kappa corrects for that:
where po is observed agreement and pe is agreement expected by chance. Worked example. Two annotators label 200 tickets as urgent or not. They agree on 170, so po=170/200=0.85. Annotator A marked 60% urgent, annotator B marked 65% urgent. Chance agreement is:
p_e = P(both say urgent) + P(both say not urgent) = (0.60 x 0.65) + (0.40 x 0.35) = 0.390 + 0.140 = 0.530kappa = (0.85 - 0.53) / (1 - 0.53) = 0.32 / 0.47 = 0.681Raw agreement of 85% sounds fine. Kappa of 0.68 says a meaningful share of that agreement was luck. Interpretation bands in common use: below 0.40 the task is poorly defined and you must rewrite the guidelines; 0.40-0.60 is weak; 0.60-0.80 is acceptable for most applications; above 0.80 is strong. If humans cannot reach 0.80 on your task, a model will not either — the ceiling is in the specification, not the model.
Inter-annotator agreement is a measurement of your task definition, not of your annotators. Low kappa means go and rewrite the guidelines.
Annotation protocols that work
- Write guidelines with worked examples, including hard cases. Ten annotated edge cases in the guidelines are worth more than three pages of prose rules.
- Run a calibration round. Everyone labels the same 50 items, you compute kappa, you discuss every disagreement, you amend the guidelines. Then start the real work. Skipping this is the most expensive shortcut in the whole pipeline.
- Double-label a 10% sample throughout so you catch drift when an annotator's interpretation quietly shifts in week three.
- Give annotators an explicit "unclear" escape hatch. Forcing a decision on genuinely ambiguous items manufactures noise. Route them to an expert instead.
- Resolve disagreements by adjudication, not majority vote, when the classes are subtle — a third annotator picking between two readings is often just a third opinion.
Crowdsourcing and active learning
Crowd platforms make sense for tasks a careful non-expert can do after reading a page of instructions: sentiment, topic, obvious-spam detection. They are wrong for anything requiring domain expertise — clinical coding, legal classification, security triage — where you need a small number of experts and should budget accordingly. When you do crowdsource, seed 5% of items with known-answer gold questions and drop workers who fall below about 90% on them.
Active learning is the technique for getting the most out of a limited labelling budget. Rather than labelling randomly, label the examples the current model is least sure about:
1import numpy as np23def selection_scores(probs):4 """probs: (n_examples, n_classes) predicted probabilities."""5 # 1. Entropy - overall uncertainty across all classes.6 entropy = -(probs * np.log(probs + 1e-12)).sum(axis=1)78 # 2. Margin - gap between top two classes. Small gap = confusable.9 top2 = np.sort(probs, axis=1)[:, -2:]10 margin = top2[:, 1] - top2[:, 0]1112 return entropy, margin1314# Label the highest-entropy / smallest-margin examples first.15# Then retrain and repeat. Each round targets the current weak spot.In practice active learning reaches a given accuracy with roughly half to two-thirds the labels of random sampling. Keep a random holdout too — a purely uncertainty-driven set is biased towards hard cases and makes a poor evaluation set.
Quality assurance before you train
| Problem | How it shows up in the model | Fix |
|---|---|---|
| Exact or near duplicates | Memorisation; inflated eval scores if duplicates cross the split | Hash exact matches; embed and drop pairs above ~0.95 cosine similarity |
| Train/test leakage | Great metrics, poor production behaviour | Split by entity (customer, document) not by row; check overlap after splitting |
| Boilerplate in the target | Model appends the boilerplate to everything | Detect strings appearing in >20% of targets and strip them |
| Class imbalance | Rare classes never predicted | Cap the majority class; set a floor per class; report per-class recall |
| Inconsistent labels for similar inputs | Model output is unstable on borderline cases | Cluster by embedding, review clusters with mixed labels |
| Truncated examples | Model learns to stop mid-sentence | Check token length distribution; drop or shorten examples exceeding max length |
| Personal data in the text | Model regurgitates real names and card numbers | Regex plus named-entity redaction before training, not after |
Encode these as a script that runs before every training job and fails loudly:
1import json, hashlib2from collections import Counter34def validate(path, tokenizer, max_len=1024, min_per_class=100):5 rows, hashes, issues = [], set(), []6 lengths, labels = [], Counter()78 for i, line in enumerate(open(path)):9 r = json.loads(line)10 if not r.get("input", "").strip() or not r.get("output", "").strip():11 issues.append(f"row {i}: empty input or output")12 continue1314 h = hashlib.md5((r["input"] + r["output"]).encode()).hexdigest()15 if h in hashes:16 issues.append(f"row {i}: exact duplicate")17 continue18 hashes.add(h)1920 n = len(tokenizer(r["input"] + r["output"])["input_ids"])21 if n > max_len:22 issues.append(f"row {i}: {n} tokens exceeds {max_len}")23 lengths.append(n)24 labels[r["output"][:40]] += 125 rows.append(r)2627 for label, count in labels.items():28 if count < min_per_class:29 issues.append(f"label '{label}': only {count} examples")3031 lengths.sort()32 print(f"kept {len(rows)} rows; token length p50={lengths[len(lengths)//2]}, "33 f"p95={lengths[int(len(lengths) * 0.95)]}, max={lengths[-1]}")34 return rows, issuesFormat and splitting
JSONL — one JSON object per line — is the practical default. It streams, it appends, a corrupt line costs you one example rather than the file, and every tool reads it.
{"id": "t-00417", "instruction": "Classify this support ticket.", "input": "my order 88213 said delivered on tuesday but nothings here, checked with neighbours too", "output": "delivery_issue", "meta": {"source": "zendesk", "annotator": "a3", "confidence": "high", "date": "2024-11-02"}}Keep the metadata. When evaluation shows a weak spot, provenance fields let you ask "are the failures concentrated in synthetic examples, or in annotator a3's work?" — a question you cannot answer without them.
Split into train, validation and test at roughly 80/10/10, and split by entity. If one customer filed 40 tickets, all 40 go to the same split; otherwise the model sees that customer's phrasing in training and you measure memorisation. The test set should be touched once, at the end. Every metric you look at repeatedly is a metric you are overfitting to.
Privacy and bias, handled at collection time
Redact before training, not before shipping. Once a name is in the weights it is not reliably removable, and models do memorise and reproduce rare strings from their fine-tuning data — long identifiers and account numbers are especially prone to it because they appear in few contexts.
1import re23PATTERNS = {4 "EMAIL": r"[\w.+-]+@[\w-]+\.[\w.]+",5 "PHONE": r"\+?\d[\d\s().-]{7,}\d",6 "CARD": r"\b(?:\d[ -]*?){13,16}\b",7 "IBAN": r"\b[A-Z]{2}\d{2}[A-Z0-9]{10,30}\b",8}910def redact(text):11 for tag, pat in PATTERNS.items():12 text = re.sub(pat, f"[{tag}]", text)13 return textRegexes catch structured identifiers; names and addresses need a named-entity model on top. Neither is perfect, so sample 200 redacted examples and read them by hand before training.
For bias, the practical check is counterfactual consistency: take 200 evaluation examples, swap a demographic attribute (name, pronoun, stated location) while holding everything else fixed, and compare outputs. If the approval rate for otherwise-identical loan queries drops when the applicant's name changes, that bias came from your training data and you can find it — look at the label distribution conditioned on the same attribute. Then correct the sampling.
Any statistical association present in your training data will be reproduced by the model, whether or not it was part of the task. The only defence is to look for it deliberately.
What this means when you build a dataset
Budget the work realistically. On a typical project, dataset construction is 60-70% of total effort and the training run is a single afternoon. Teams that plan the reverse ratio end up with a beautifully configured pipeline training on data nobody inspected.
Concretely, before your first real training run: read 100 random examples yourself, end to end. Not a summary, not a sample of the schema — read them. Every dataset problem described above is visible to a human reading 100 examples in twenty minutes, and none of them are visible in a loss curve. The team with the "anything else I can help you with today?" model would have caught it on example four.
Then version the dataset like code. Give each build a version string, record which examples went in, and store it alongside the model checkpoint it produced. When a model behaves strangely in three months, the first question will be "what was it trained on?" and you want that to be a lookup, not an investigation.