Course Content
Fine-Tuning LLMs
6 sections · 52 lessons
Roughly how much data do you need to train from scratch versus to fine-tune a base model?
What you need to know
Why the gap is so large
Pretraining teaches a model language, facts and reasoning from nothing, starting from random weights. Fine-tuning starts from a model that already knows all that, and only changes behaviour: the format, the task, the tone. Changing behaviour needs far less evidence than learning language.
The numbers, in tokens
| Stage | Typical data | In tokens |
|---|---|---|
| Pretraining from scratch | Web, books, code | 10–36 trillion for recent open models |
| Continual pretraining | Raw domain or language text | 100 million to tens of billions |
| Instruction fine-tuning (SFT) | 1,000–50,000 examples | About 0.5 to 25 million |
| Narrow task (classify, extract) | 500–5,000 examples | About 0.25 to 5 million |
| Preference tuning (DPO) | 2,000–20,000 pairs | A few million to tens of millions |
A worked comparison: 5,000 examples of 500 tokens is 2.5 million tokens. Llama 3 was pretrained on over 15 trillion tokens — about six million times more.
Chinchilla, and why models go past it
The 2022 Chinchilla study found that for a fixed training compute budget, about 20 tokens per parameter gives the best model. For an 8B model that is about 160 billion tokens. Modern small models are trained far beyond this — Llama 3's 8B model saw over 15 trillion tokens, nearly 1,900 per parameter — because a smaller model trained longer is cheaper to serve for its whole life. Meta reported about 39 million H100 GPU-hours to train the Llama 3.1 family. That is why almost nobody outside a few labs pretrains from scratch.
Quality over quantity
The 2023 LIMA paper built a strong chat model from about 1,000 carefully chosen examples. The lesson is not "1,000 is always enough", but that consistency and coverage matter more than volume. Contradictory labels set a ceiling that no amount of extra data can break.
Plot a learning curve
Train on 25%, 50% and 100% of your data and evaluate each on the same held-out set:
- Still rising steeply — more data of the same kind will help.
- Flat — more of the same will not help. Fix label quality, cover missing cases, or rethink the task.
Hosted APIs accept very small sets (OpenAI's minimum is 10 examples, and its guidance suggests improvements often show from 50–100), which is fine for format and tone, not for hard tasks.
A real-life example
A Mumbai law firm fine-tunes a legal-clause classifier for 14 clause types in Indian commercial contracts. It has 6,000 labelled clauses and wonders whether to pay paralegals to label 6,000 more.
The learning curve on 800 held-out clauses:
| Training clauses | Macro-F1 |
|---|---|
| 1,500 | 0.78 |
| 3,000 | 0.85 |
| 6,000 | 0.87 |
The curve is flattening, so another 6,000 of the same would add perhaps a point. The confusion matrix shows most errors between "indemnity" and "limitation of liability", and a review finds three paralegals labelled mixed clauses differently. The firm writes one clear rule for mixed clauses and relabels 400 examples. Macro-F1 goes to 0.90 — more than the extra 6,000 labels would have given, for a fraction of the cost. (Made-up numbers for illustration.)
Follow-up questions to expect
- "When do you need much more data than a few thousand examples?" — When you are teaching new knowledge or a new language, not behaviour. Then continual pretraining on millions of tokens comes before SFT.
- "How many examples per class for a classifier?" — A common starting point is a few hundred per class, with more for classes that are easily confused. Let the learning curve decide.
- "Can synthetic data fill the gap?" — Yes, if it is filtered and checked against a human-labelled eval set; unchecked synthetic data can teach the generator's mistakes.