LLMs Deep Dive

Course Content

LLMs Deep Dive

10 sections · 40 lessons

What is overfitting, and how can it be prevented?


Fine-tuning on 2,000 support chats1.201.250.851.100.501.150.251.35Training lossValidation lossEpoch 1Epoch 2Epoch 3Epoch 4Early stopping keeps the epoch-2 checkpoint.
After epoch 2 the model is learning the chats, not the job — training loss keeps falling while validation loss turns up.

What you need to know

Reading the loss curves

A fine-tuning run on 2,000 support chats:

EpochTraining lossValidation loss
11.201.25
20.851.10
30.501.15
40.251.35

Epoch 2 is the best checkpoint. After it, training loss keeps falling but validation loss rises: the model is memorising specific chats rather than learning how to reply. Early stopping keeps the epoch-2 weights.

Prevention toolkit

  • More, more varied data — or augmentation (paraphrases, different languages) when collection is hard.
  • Early stopping — keep the checkpoint with the best validation score.
  • Regularisation — weight decay (L2), dropout, label smoothing.
  • Less capacity — a smaller model, or LoRA with a low rank instead of full fine-tuning.
  • Clean splits — remove duplicates and near-duplicates across train, validation and test; the same chat appearing in both gives fake good scores.
  • Cross-validation — for small datasets, so one lucky split does not drive decisions.

How overfitting looks in LLMs

  • Memorisation — the model repeats training text word for word, including names, phone numbers or account details. This is a privacy risk as well as a quality problem.
  • Style collapse — every reply starts with the same phrase from the training data.
  • Brittleness — works on phrasings seen in training, fails on small rewordings.
  • Benchmark contamination — test questions leaked into pretraining data make a model look better than it is. Use fresh, private evaluation sets.

Large pretrained models rarely overfit during pretraining, because they see most data only once or a few times. The risk is in fine-tuning on small datasets.

A real-life example

A bank fine-tunes a model on 2,000 chats for 6 epochs. In testing it looks excellent on familiar questions. Then a reviewer notices that for a question about a failed transfer, the bot replied "Dear Mr. Sharma, your transfer of Rs 12,500 to account ending 4471..." — details from one training chat, repeated to a different customer.

The team takes three steps. They mask names, account numbers and amounts in the training data. They switch to LoRA, 2 epochs, with early stopping on a validation set of 300 chats from a different month. And they add a memorisation test: prompt the model with the first half of 100 training chats and check how often it completes the second half word for word. Replies are now more general, and no training-set personal data appears in outputs.

Follow-up questions to expect

  • "How do you tell overfitting from underfitting?" — Overfitting: low training loss, higher validation loss. Underfitting: both high. The fix for one makes the other worse.
  • "Why does dropout help?" — It randomly switches off units during training, so the network cannot depend on any single path and learns more robust features.
  • "Can a large model overfit a small dataset?" — Very easily; it has enough capacity to memorise thousands of examples in a few epochs.