Course Content
Fine-Tuning LLMs
6 sections · 52 lessons
What is synthetic data, and how can you generate and use it for fine-tuning safely?
What you need to know
Common patterns
| Pattern | What the generator does | Good for |
|---|---|---|
| Answer generation | Writes answers for real user prompts | Bulk SFT data |
| Seed-and-expand (Self-Instruct style) | Writes new prompts from a few seed examples | Covering rare intents |
| Back-generation | Writes the questions a document answers | RAG and domain Q&A |
| Rejection sampling | Samples many answers; a checker keeps the correct ones | Maths, code, reasoning |
| Preference pairs | Samples two answers; a judge picks the better one | DPO data |
What goes wrong
- Copied errors. The student learns the teacher's mistakes, stated confidently.
- Low diversity. Generators repeat themselves: the same openings, the same structure. The student becomes narrow.
- Distribution drift. Invented prompts are cleaner and more polite than real user messages.
- Model collapse. A 2024 Nature paper (Shumailov et al.) showed that models trained again and again on their own outputs lose the rare "tail" cases of the original data.
- Licences. Some API providers' terms forbid using outputs to build competing models (see the distillation lesson).
A safe pipeline
- Anchor on reality — real prompts where possible; if you must invent them, seed from real examples.
- Generate with a spec — give the generator your written spec and gold examples.
- Verify — schema checks, run the code, check facts against source documents, and score with an LLM judge.
- Deduplicate and check diversity — near-duplicate removal; look at clusters of openings and topics.
- Human spot check — read a random 2–5% sample; if the error rate is high, fix the generator, not just the rows.
- Evaluate on real data — the test set must be real, human-labelled and never generated by the same model.
1from pydantic import BaseModel, ValidationError23class Reply(BaseModel):4 intent: str5 reply_hinglish: str6 needs_escalation: bool78def keep(raw: str, allowed_intents: set[str]) -> bool:9 try:10 r = Reply.model_validate_json(raw)11 except ValidationError:12 return False # broken JSON or missing field13 return r.intent in allowed_intents and 5 <= len(r.reply_hinglish.split()) <= 80Cheap rule checks like this remove the worst rows before you pay for an LLM judge on the rest.
A real-life example
The Hindi support team has only 800 real chats about rare issues, such as refunding a cash-on-delivery order to a bank account. They write 200 seed messages and a list of 40 intents, and ask a strong model to write 6,000 new customer messages in realistic Hinglish, with typos and mixed scripts, plus replies grounded in the current policy pages.
Filtering removes 30% as near-duplicates and 8% that quote a wrong refund timeline. Agents review 300 random rows and find 4% still wrong, which the team accepts. The test set stays 300 real chats. The fine-tune improves rare-intent accuracy on real chats, not only on synthetic ones — the result that matters.
Follow-up questions to expect
- "How much synthetic data should you mix with real data?" — There is no fixed ratio. Keep all good real data, add synthetic data where error analysis shows gaps, and check that real-data scores improve.
- "How do you spot low diversity?" — Count distinct first sentences, cluster embeddings and look for a few giant clusters, and compare length and vocabulary with real data.
- "Can you use the same model as generator and judge?" — It is risky: it tends to approve its own style and miss its own errors. Use a different model or human checks for the judge.