Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

What are the major stages in an end-to-end pipeline for building a strong LLM?


From raw text to a deployed assistantPretraining: trillions of tokensMid-training: code, long contextSFT: instructions, chats, tool callsPreferences (DPO), then RL with checkersEvaluation, red-teaming, quantisation
Each stage down is cheaper and faster, which is why product teams almost always start at the third one.

What you need to know

StageDataObjectiveScale
1. PretrainingTrillions of tokens of filtered web, books, codeNext-token predictionAlmost all the compute
2. Mid-training / continued pretrainingTargeted text: long documents, code, maths, a languageNext-token predictionBillions of tokens
3. Supervised fine-tuningInstructions, multi-turn chats, tool-call tracesImitate ideal answersThousands to millions of examples
4. Preference optimisationChosen vs rejected pairsDPO, or reward model + RLThousands to millions of pairs
5. RL with verifiable rewardsProblems with checkable answersGRPO and similarMany generated attempts per problem
6. Evaluation and deploymentBenchmarks, red-team prompts, live trafficMeasure, harden, serveOngoing

What each stage gives the model

  • Pretraining gives knowledge and general skill. The output is a base model that completes text but does not reliably follow instructions.
  • SFT gives the chat template, instruction following and task formats.
  • Preference optimisation gives the negative signal SFT lacks: which answers to avoid — unhelpful, unsafe, sycophantic.
  • RL with verifiable rewards became central after reasoning models such as OpenAI's o1 (2024) and DeepSeek-R1 (2025). The model writes many attempts; a checker (correct final answer, passing unit tests) rewards the good ones; GRPO updates the model. Long step-by-step reasoning emerges from this stage.
  • Deployment work includes safety testing, quantisation or distillation into smaller models, guardrails and monitoring.

Where product teams fit

Almost no product team pretrains. They start from an open-weight base or instruct model and work in stages 3–6: SFT on their task, perhaps DPO on their preferences, perhaps GRPO if the task has a checker, then evaluation and serving. Each stage down the table is cheaper and faster, and each is limited by data quality more than by the choice of optimizer.

A real-life example

An Indian startup builds a Hindi-first assistant from an open-weight 8B base model:

  1. Continued pretraining on 20 billion tokens of Hindi and other Indian-language text, with 25% English replay.
  2. SFT on 200,000 instruction and chat examples, half translated and reviewed, half written by native speakers.
  3. DPO on 30,000 preference pairs labelled against a rubric that includes "matches the user's script: Devanagari or romanised Hindi".
  4. GRPO on 15,000 school-level maths problems written in Hindi, rewarded by exact final answers.
  5. Evaluation on Hindi versions of reasoning and safety tests, plus red-teaming by native speakers; then 4-bit quantisation for serving.

Each stage has its own dataset, its own eval, and a clear question it answers.

Follow-up questions to expect

  • "Where does most of the quality come from?" — Knowledge and general skill from pretraining; usefulness, style and reasoning from post-training. A weak base cannot be rescued by post-training.
  • "Why do SFT before preference training?" — Preference methods assume the model already produces reasonable answers to compare.
  • "What is mid-training?" — A late pretraining phase on a curated mix (long documents, code, maths) that prepares the model for post-training without starting over.