Course Content
Fine-Tuning LLMs
6 sections · 52 lessons
What are the major stages in an end-to-end pipeline for building a strong LLM?
What you need to know
| Stage | Data | Objective | Scale |
|---|---|---|---|
| 1. Pretraining | Trillions of tokens of filtered web, books, code | Next-token prediction | Almost all the compute |
| 2. Mid-training / continued pretraining | Targeted text: long documents, code, maths, a language | Next-token prediction | Billions of tokens |
| 3. Supervised fine-tuning | Instructions, multi-turn chats, tool-call traces | Imitate ideal answers | Thousands to millions of examples |
| 4. Preference optimisation | Chosen vs rejected pairs | DPO, or reward model + RL | Thousands to millions of pairs |
| 5. RL with verifiable rewards | Problems with checkable answers | GRPO and similar | Many generated attempts per problem |
| 6. Evaluation and deployment | Benchmarks, red-team prompts, live traffic | Measure, harden, serve | Ongoing |
What each stage gives the model
- Pretraining gives knowledge and general skill. The output is a base model that completes text but does not reliably follow instructions.
- SFT gives the chat template, instruction following and task formats.
- Preference optimisation gives the negative signal SFT lacks: which answers to avoid — unhelpful, unsafe, sycophantic.
- RL with verifiable rewards became central after reasoning models such as OpenAI's o1 (2024) and DeepSeek-R1 (2025). The model writes many attempts; a checker (correct final answer, passing unit tests) rewards the good ones; GRPO updates the model. Long step-by-step reasoning emerges from this stage.
- Deployment work includes safety testing, quantisation or distillation into smaller models, guardrails and monitoring.
Where product teams fit
Almost no product team pretrains. They start from an open-weight base or instruct model and work in stages 3–6: SFT on their task, perhaps DPO on their preferences, perhaps GRPO if the task has a checker, then evaluation and serving. Each stage down the table is cheaper and faster, and each is limited by data quality more than by the choice of optimizer.
A real-life example
An Indian startup builds a Hindi-first assistant from an open-weight 8B base model:
- Continued pretraining on 20 billion tokens of Hindi and other Indian-language text, with 25% English replay.
- SFT on 200,000 instruction and chat examples, half translated and reviewed, half written by native speakers.
- DPO on 30,000 preference pairs labelled against a rubric that includes "matches the user's script: Devanagari or romanised Hindi".
- GRPO on 15,000 school-level maths problems written in Hindi, rewarded by exact final answers.
- Evaluation on Hindi versions of reasoning and safety tests, plus red-teaming by native speakers; then 4-bit quantisation for serving.
Each stage has its own dataset, its own eval, and a clear question it answers.
Follow-up questions to expect
- "Where does most of the quality come from?" — Knowledge and general skill from pretraining; usefulness, style and reasoning from post-training. A weak base cannot be rescued by post-training.
- "Why do SFT before preference training?" — Preference methods assume the model already produces reasonable answers to compare.
- "What is mid-training?" — A late pretraining phase on a curated mix (long documents, code, maths) that prepares the model for post-training without starting over.