Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

Roughly how much data do you need to train from scratch versus to fine-tune a base model?


Macro-F1 of the clause classifier as data grows0.780.850.870.9001231,500 clauses3,0006,0006,000, 400relabelledIllustrative numbers from the lesson's made-up scenario.
A flattening curve says more of the same data will not help — fixing 400 inconsistent labels beat doubling the set.

What you need to know

Why the gap is so large

Pretraining teaches a model language, facts and reasoning from nothing, starting from random weights. Fine-tuning starts from a model that already knows all that, and only changes behaviour: the format, the task, the tone. Changing behaviour needs far less evidence than learning language.

The numbers, in tokens

StageTypical dataIn tokens
Pretraining from scratchWeb, books, code10–36 trillion for recent open models
Continual pretrainingRaw domain or language text100 million to tens of billions
Instruction fine-tuning (SFT)1,000–50,000 examplesAbout 0.5 to 25 million
Narrow task (classify, extract)500–5,000 examplesAbout 0.25 to 5 million
Preference tuning (DPO)2,000–20,000 pairsA few million to tens of millions

A worked comparison: 5,000 examples of 500 tokens is 2.5 million tokens. Llama 3 was pretrained on over 15 trillion tokens — about six million times more.

Chinchilla, and why models go past it

The 2022 Chinchilla study found that for a fixed training compute budget, about 20 tokens per parameter gives the best model. For an 8B model that is about 160 billion tokens. Modern small models are trained far beyond this — Llama 3's 8B model saw over 15 trillion tokens, nearly 1,900 per parameter — because a smaller model trained longer is cheaper to serve for its whole life. Meta reported about 39 million H100 GPU-hours to train the Llama 3.1 family. That is why almost nobody outside a few labs pretrains from scratch.

Quality over quantity

The 2023 LIMA paper built a strong chat model from about 1,000 carefully chosen examples. The lesson is not "1,000 is always enough", but that consistency and coverage matter more than volume. Contradictory labels set a ceiling that no amount of extra data can break.

Plot a learning curve

Train on 25%, 50% and 100% of your data and evaluate each on the same held-out set:

  • Still rising steeply — more data of the same kind will help.
  • Flat — more of the same will not help. Fix label quality, cover missing cases, or rethink the task.

Hosted APIs accept very small sets (OpenAI's minimum is 10 examples, and its guidance suggests improvements often show from 50–100), which is fine for format and tone, not for hard tasks.

A real-life example

A Mumbai law firm fine-tunes a legal-clause classifier for 14 clause types in Indian commercial contracts. It has 6,000 labelled clauses and wonders whether to pay paralegals to label 6,000 more.

The learning curve on 800 held-out clauses:

Training clausesMacro-F1
1,5000.78
3,0000.85
6,0000.87

The curve is flattening, so another 6,000 of the same would add perhaps a point. The confusion matrix shows most errors between "indemnity" and "limitation of liability", and a review finds three paralegals labelled mixed clauses differently. The firm writes one clear rule for mixed clauses and relabels 400 examples. Macro-F1 goes to 0.90 — more than the extra 6,000 labels would have given, for a fraction of the cost. (Made-up numbers for illustration.)

Follow-up questions to expect

  • "When do you need much more data than a few thousand examples?" — When you are teaching new knowledge or a new language, not behaviour. Then continual pretraining on millions of tokens comes before SFT.
  • "How many examples per class for a classifier?" — A common starting point is a few hundred per class, with more for classes that are easily confused. Let the learning curve decide.
  • "Can synthetic data fill the gap?" — Yes, if it is filtered and checked against a human-labelled eval set; unchecked synthetic data can teach the generator's mistakes.