Course Content
Fine-Tuning LLMs
6 sections · 52 lessons
If a fine-tuned model is wrong because the training data is noisy, how do you diagnose and fix it?
What you need to know
Diagnose
- Error taxonomy — read 100–200 failures and name the error types. "It invents refund timelines in 30% of refund questions" is actionable; "it is wrong" is not.
- Trace to training rows — embed each failing prompt and find its nearest training examples. Look at what they teach.
- Rank rows by loss — very high loss after training often means the label is wrong or contradicts other rows.
- Find contradictions — group near-duplicate prompts and flag groups with different answers.
- Split by source — compute error rates for each data source (vendor, generator, team). One source is often the problem.
1import torch23@torch.no_grad()4def row_loss(model, tok, messages):5 text = tok.apply_chat_template(messages, tokenize=False)6 ids = tok(text, return_tensors="pt").input_ids.to(model.device)7 return model(input_ids=ids, labels=ids).loss.item()89# losses = [row_loss(model, tok, r["messages"]) for r in train_rows]10# then read the 50 highest-loss rows by handThis scores whole conversations, prompt included, which is fine for ranking. For classification labels, "confident learning" tools such as cleanlab flag rows where a trained model strongly disagrees with the given label.
Fix
- Fix the spec first if labellers disagreed because the rule was unclear; otherwise relabelling reproduces the same disagreement.
- Relabel or remove the bad cluster, then retrain. This alone often recovers most of the gap.
- Add a filter to the data pipeline (rule check or LLM judge) so the same problem cannot return.
- If you must keep noisy data, train on the verified subset first and add the rest with a lower weight.
Verify
Check on a freshly labelled test set. If the old test set came from the same noisy source, it will tell you the problem is solved when it is not.
A real-life example
The Hindi support model tells some customers refunds take "7 working days". The policy says 5.
The team's error taxonomy shows 28% of refund answers have wrong timelines. Nearest-neighbour search on those prompts finds that the closest training examples all came from one synthetic batch, generated from an old policy page. The loss ranking shows no alarm — the wrong rows were consistent with each other, so the model learned them easily, which is why loss alone did not catch this.
They delete the 1,100 rows from that batch, regenerate them from the current policy, add a rule that checks every stated timeline against the policy table, and retrain. On 200 freshly labelled refund chats, wrong timelines drop to under 2%.
Follow-up questions to expect
- "Why not just train longer or use a bigger model?" — A model trained on contradictions learns contradictions. More capacity fits the noise better.
- "How do you find mislabelled rows in a million-row dataset?" — Loss ranking, confident-learning tools, and an LLM judge with the spec, then human review of the flagged rows only.
- "How do you stop this next time?" — Record the source and version of every row, and run the automatic checks in the data pipeline, not after training.