Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

What is synthetic data, and how can you generate and use it for fine-tuning safely?


What you need to know

Common patterns

PatternWhat the generator doesGood for
Answer generationWrites answers for real user promptsBulk SFT data
Seed-and-expand (Self-Instruct style)Writes new prompts from a few seed examplesCovering rare intents
Back-generationWrites the questions a document answersRAG and domain Q&A
Rejection samplingSamples many answers; a checker keeps the correct onesMaths, code, reasoning
Preference pairsSamples two answers; a judge picks the better oneDPO data

What goes wrong

  • Copied errors. The student learns the teacher's mistakes, stated confidently.
  • Low diversity. Generators repeat themselves: the same openings, the same structure. The student becomes narrow.
  • Distribution drift. Invented prompts are cleaner and more polite than real user messages.
  • Model collapse. A 2024 Nature paper (Shumailov et al.) showed that models trained again and again on their own outputs lose the rare "tail" cases of the original data.
  • Licences. Some API providers' terms forbid using outputs to build competing models (see the distillation lesson).

A safe pipeline

  1. Anchor on reality — real prompts where possible; if you must invent them, seed from real examples.
  2. Generate with a spec — give the generator your written spec and gold examples.
  3. Verify — schema checks, run the code, check facts against source documents, and score with an LLM judge.
  4. Deduplicate and check diversity — near-duplicate removal; look at clusters of openings and topics.
  5. Human spot check — read a random 2–5% sample; if the error rate is high, fix the generator, not just the rows.
  6. Evaluate on real data — the test set must be real, human-labelled and never generated by the same model.
Python
from pydantic import BaseModel, ValidationErrorclass Reply(BaseModel):    intent: str    reply_hinglish: str    needs_escalation: booldef keep(raw: str, allowed_intents: set[str]) -> bool:    try:        r = Reply.model_validate_json(raw)    except ValidationError:        return False                     # broken JSON or missing field    return r.intent in allowed_intents and 5 <= len(r.reply_hinglish.split()) <= 80

Cheap rule checks like this remove the worst rows before you pay for an LLM judge on the rest.

A real-life example

The Hindi support team has only 800 real chats about rare issues, such as refunding a cash-on-delivery order to a bank account. They write 200 seed messages and a list of 40 intents, and ask a strong model to write 6,000 new customer messages in realistic Hinglish, with typos and mixed scripts, plus replies grounded in the current policy pages.

Filtering removes 30% as near-duplicates and 8% that quote a wrong refund timeline. Agents review 300 random rows and find 4% still wrong, which the team accepts. The test set stays 300 real chats. The fine-tune improves rare-intent accuracy on real chats, not only on synthetic ones — the result that matters.

Follow-up questions to expect

  • "How much synthetic data should you mix with real data?" — There is no fixed ratio. Keep all good real data, add synthetic data where error analysis shows gaps, and check that real-data scores improve.
  • "How do you spot low diversity?" — Count distinct first sentences, cluster embeddings and look for a few giant clusters, and compare length and vocabulary with real data.
  • "Can you use the same model as generator and judge?" — It is risky: it tends to approve its own style and miss its own errors. Use a different model or human checks for the judge.