Course Content
Fine-Tuning LLMs
6 sections · 52 lessons
How do you reduce overfitting and memorization when a model repeats training data verbatim?
What you need to know
Why it happens
- Repetition. Research on memorisation (Carlini et al., 2022) found that the more often a sequence appears in training, the more likely it is memorised. Duplicates are the biggest driver.
- Too many epochs. On a small dataset, each extra pass pushes the model from learning the pattern toward storing the examples.
- Too much capacity for the data. A high LoRA rank or full fine-tuning on 1,000 examples can store them.
How to measure it
1def ngrams(tokens, n=20):2 return {tuple(tokens[i:i + n]) for i in range(len(tokens) - n + 1)}34train_grams = set()5for doc in train_docs:6 train_grams |= ngrams(doc.split())78def copies_training(output: str) -> bool:9 return bool(ngrams(output.split()) & train_grams) # any 20-word overlapRun this over a few hundred generated outputs and track the share that contains a 20-word span from the training data. Also watch the loss curves: validation loss rising while training loss keeps falling is the classic overfitting sign.
Fixes, cheapest first
- Deduplicate exact and near-duplicate examples (MinHash or embedding similarity).
- Fewer epochs, early stopping — stop when validation loss stops improving; 1–3 epochs is normal.
- Lower the learning rate or LoRA rank.
- Regularise —
lora_dropoutof 0.05–0.1, weight decay, and optionally NEFTune (random noise added to embeddings in training;neftune_noise_alphain Hugging FaceTrainingArguments). - More varied data — paraphrases and more sources.
- Mask the prompt so loss is only on answers, and the model does not memorise boilerplate prompts.
For sensitive data, removing it from training is the only real fix. Differential-privacy training (DP-SGD) gives formal guarantees, at a noticeable cost in quality.
A real-life example
A radiology summariser starts writing, in a new patient's summary, a full sentence that includes another patient's name from the training data. That is a privacy incident, not just a quality bug.
The overlap check shows 3% of outputs contain a 20-word span copied from training. The causes: 4 epochs on only 1,800 examples, and templated normal reports that appeared about 30 times each, some with names left in the free-text section. The team scrubs names with a PHI detector plus manual review, deduplicates, trains 2 epochs with early stopping, and adds an output filter that blocks any person's name not present in the input report. The copied-span rate falls to 0.2%, and none of those spans contain personal data.
Follow-up questions to expect
- "Is some memorisation acceptable?" — Yes for standard phrases, legal boilerplate or a required disclaimer. The problem is copying private or unique content.
- "How would an attacker extract memorised data?" — By prompting with the beginning of a likely training record and letting the model complete it. Test your model the same way before release.
- "Does LoRA prevent memorisation?" — It limits capacity, which helps, but a LoRA adapter trained for many epochs on duplicates will still memorise.