Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

How do you reduce overfitting and memorization when a model repeats training data verbatim?


What you need to know

Why it happens

  • Repetition. Research on memorisation (Carlini et al., 2022) found that the more often a sequence appears in training, the more likely it is memorised. Duplicates are the biggest driver.
  • Too many epochs. On a small dataset, each extra pass pushes the model from learning the pattern toward storing the examples.
  • Too much capacity for the data. A high LoRA rank or full fine-tuning on 1,000 examples can store them.

How to measure it

Python
def ngrams(tokens, n=20):    return {tuple(tokens[i:i + n]) for i in range(len(tokens) - n + 1)}train_grams = set()for doc in train_docs:    train_grams |= ngrams(doc.split())def copies_training(output: str) -> bool:    return bool(ngrams(output.split()) & train_grams)   # any 20-word overlap

Run this over a few hundred generated outputs and track the share that contains a 20-word span from the training data. Also watch the loss curves: validation loss rising while training loss keeps falling is the classic overfitting sign.

Fixes, cheapest first

  1. Deduplicate exact and near-duplicate examples (MinHash or embedding similarity).
  2. Fewer epochs, early stopping — stop when validation loss stops improving; 1–3 epochs is normal.
  3. Lower the learning rate or LoRA rank.
  4. Regularise — lora_dropout of 0.05–0.1, weight decay, and optionally NEFTune (random noise added to embeddings in training; neftune_noise_alpha in Hugging Face TrainingArguments).
  5. More varied data — paraphrases and more sources.
  6. Mask the prompt so loss is only on answers, and the model does not memorise boilerplate prompts.

For sensitive data, removing it from training is the only real fix. Differential-privacy training (DP-SGD) gives formal guarantees, at a noticeable cost in quality.

A real-life example

A radiology summariser starts writing, in a new patient's summary, a full sentence that includes another patient's name from the training data. That is a privacy incident, not just a quality bug.

The overlap check shows 3% of outputs contain a 20-word span copied from training. The causes: 4 epochs on only 1,800 examples, and templated normal reports that appeared about 30 times each, some with names left in the free-text section. The team scrubs names with a PHI detector plus manual review, deduplicates, trains 2 epochs with early stopping, and adds an output filter that blocks any person's name not present in the input report. The copied-span rate falls to 0.2%, and none of those spans contain personal data.

Follow-up questions to expect

  • "Is some memorisation acceptable?" — Yes for standard phrases, legal boilerplate or a required disclaimer. The problem is copying private or unique content.
  • "How would an attacker extract memorised data?" — By prompting with the beginning of a likely training record and letting the model complete it. Test your model the same way before release.
  • "Does LoRA prevent memorisation?" — It limits capacity, which helps, but a LoRA adapter trained for many epochs on duplicates will still memorise.