Course Content
How Large Language Models Work
3 sections · 9 lessons
Fine-Tuning vs Pretraining vs RLHF
Download a raw pretrained model — one that has been trained on text and nothing else — and ask it a question:
Prompt: What is the capital of France?Output: What is the largest country in Europe? What is the longest river in the world? What is the highest mountain? Answers to these and other questions can be found in our quiz pack...The model is not broken. It is working perfectly. It was trained to continue text, and on the open internet, a line that looks like a quiz question is most often followed by more quiz questions. It gave the statistically best continuation. It just was not the one you wanted.
This gap — between a model that predicts text and a model that answers you — is what the training pipeline after pretraining exists to close. The classic pipeline has three stages, with three completely different data types, three cost profiles differing by four orders of magnitude. Confusing them is the single most common source of bad decisions about how to adapt a model, so it is worth being precise about what each one actually changes.
The three stages at a glance
| Pretraining | Supervised fine-tuning | Preference optimisation | |
|---|---|---|---|
| Question it answers | How does language work? | What shape should a reply have? | Which of two valid replies is better? |
| Data | Trillions of tokens of raw text | 10k–1M prompt/response pairs | 10k–1M ranked comparisons |
| Labels come from | The text itself | Humans writing or curating answers | Humans choosing between outputs |
| Compute | 1021–1025 FLOPs | 1017–1019 FLOPs | 1018–1020 FLOPs |
| Wall-clock | Weeks to months on thousands of accelerators | Hours to days on 8 | Days, plus months of human labelling |
| What it adds | Knowledge and capability | Format and instruction-following | Judgement about quality, tone, safety |
| Fraction of total compute | ~99% | <1% | <1% |
Almost all of a model's ability is created in pretraining. Everything after it is teaching the model to use abilities it already has — which is why it can be done with a rounding error's worth of compute.
That last row describes the classic chat-model pipeline. Reasoning models added a fourth stage — reinforcement learning on problems whose answers can be checked, covered below — and it can absorb a much larger share of compute than the <1% shown here, because it is limited by how many problems you can generate and check, not by human labelling.
Pretraining: where the capability comes from
The objective is the plainest one available: predict the next token, and take the loss as −logp of the true token. No labels, no annotation, no human in the loop. Any text is training data, which is what makes the scale possible.
The compute required follows a reliable approximation:
with N parameters and D training tokens. The 6 comes from roughly 2 FLOPs per parameter per token in the forward pass and about 4 in the backward pass. Work a 7-billion-parameter model trained on 2 trillion tokens:
C = 6 x 7e9 x 2e12 = 8.4e22 FLOPsAt an effective 1.2e14 FLOP/s per accelerator (about 40% of an A100's bf16 peak): 8.4e22 / 1.2e14 = 7.0e8 seconds = 8,100 accelerator-daysOn 1,024 accelerators: about 8 days of continuous running.Add data collection, deduplication, filtering, failed runs and evaluation, and the realistic cost of a competent 7B model is measured in millions of dollars and a team of experienced people. This is the number that determines who gets to pretrain and who does not.
A crucial constraint shapes the design. Compute-optimal scaling work found that, for a fixed budget, model size and data should grow together at roughly 20 tokens per parameter. Earlier large models were badly undertrained on this measure — parameters were scaled aggressively while data lagged — and a smaller model trained on more tokens beat them. Models intended for heavy deployment go further still, training well past the compute-optimal point, because a smaller model costs less on every single inference call thereafter.
What exists after pretraining, and what does not
| Present | Absent |
|---|---|
| Grammar, syntax, discourse structure | Any notion that a question should be answered |
| Broad world knowledge | A consistent persona or voice |
| Multiple languages and programming languages | Refusal of harmful requests |
| Arithmetic and code semantics | Knowing when to say "I don't know" |
| In-context learning from examples in the prompt | Any preference for helpful over merely plausible |
The right-hand column is not a list of missing knowledge. It is a list of missing behaviours. That distinction is the key to everything that follows.
Supervised fine-tuning: teaching the shape of a reply
SFT continues the exact same next-token objective, on a different dataset: prompt–response pairs written or curated by people.
1{2 "messages": [3 {"role": "user", "content": "Why does bread rise?"},4 {"role": "assistant", "content": "Yeast in the dough consumes sugars and releases carbon dioxide. The gluten network formed when the dough is kneaded traps those gas bubbles, and the trapped gas expands the dough."}5 ]6}One detail separates SFT from ordinary training, and getting it wrong quietly ruins the result: loss masking. The loss is computed only on the assistant's tokens. The user's tokens are set to the ignore index and contribute nothing.
1import torch23IGNORE = -10045def build_example(tokenizer, prompt, response):6 p_ids = tokenizer(prompt, add_special_tokens=False).input_ids7 r_ids = tokenizer(response, add_special_tokens=False).input_ids89 input_ids = p_ids + r_ids10 labels = [IGNORE] * len(p_ids) + r_ids # only the response is learned1112 return torch.tensor(input_ids), torch.tensor(labels)Skip the masking and you are training the model to generate user questions as well as answers — which is exactly the failure the whole stage exists to fix. Models fine-tuned without masking often start producing a plausible answer and then inventing the user's follow-up question.
Quality dominates quantity here
This is the finding that most surprises people coming from a pretraining mindset. Careful work has shown that a few thousand meticulously written, diverse examples can produce better instruction-following than hundreds of thousands of noisy, machine-generated ones. The reason follows from the table above: the model already possesses the capability. SFT is demonstrating a format, and a thousand clear demonstrations beat a hundred thousand muddled ones.
The practical implication is uncomfortable but useful: if your fine-tune is disappointing, adding more data of the same quality is usually the wrong response. Read fifty of your own examples carefully. The problem is normally in there.
Preference optimisation: choosing between valid answers
SFT teaches the model to produce an answer. It cannot teach it which of several correct answers is better, because the training signal only ever shows one target. Consider two responses to "my code throws a null pointer exception":
- "A null pointer exception occurs when you dereference a null reference." — true, and useless.
- "Something is null when you use it. Check the variable on the line the trace points to, and trace back to where it was last assigned — a function returning null on a failure path is the usual cause." — also true, and actually helpful.
Both are valid continuations. Nobody can easily write a loss function that prefers the second. But almost anyone can look at the pair and pick. That asymmetry — hard to specify, easy to judge — is what preference optimisation exploits.
The reward-model route (RLHF)
Three steps.
Collect comparisons. Sample several responses from the SFT model for each prompt, and have annotators rank them. The output is triples: prompt x, preferred response yw, rejected response yl.
Train a reward model. Take a copy of the model, replace the vocabulary-sized output head with a single scalar, and train it so preferred responses score higher. The loss comes from the Bradley–Terry model of pairwise choice:
Only the difference in scores appears, so the reward model never needs an absolute scale — it only needs to get orderings right.
Optimise the policy. Use reinforcement learning (classically PPO) to maximise reward, with a penalty that keeps the model from straying far from where it started:
objective=E[rϕ(x,y)]−βKL(πθ∥πref)The KL term is not optional. Without it the policy discovers that the reward model is an imperfect proxy and games it — a phenomenon called reward hacking. Typical symptoms are dryly funny: outputs balloon in length because annotators mildly preferred longer answers; every reply opens with effusive agreement because that scored well; the model produces confident-sounding structure with no content. The reward went up. The model got worse.
A reward model is a learned approximation of human judgement, and optimising hard against any approximation eventually optimises the gap rather than the goal.
The direct route (DPO)
The RLHF pipeline is heavy: three models in memory at once, an RL loop that is notoriously fiddly to stabilise. Direct Preference Optimisation showed that the reward model can be skipped entirely. The same optimum is reachable with a plain supervised loss on the preference pairs:
LDPO=−logσ(βlogπref(yw∣x)πθ(yw∣x)−βlogπref(yl∣x)πθ(yl∣x))Read it as: raise the model's probability of the preferred response relative to the frozen reference, lower it for the rejected one, and let β control how far from the reference the model may drift. No reward model, no sampling loop, no RL.
RLHF (PPO) DPO Models needed in memory Policy, reference, reward, value Policy and reference Training loop RL — sensitive to hyperparameters Supervised — stable Can score new samples on the fly Yes, via the reward model No — limited to the collected pairs Typical use Large training runs with continuous data collection Much open-model alignment work, often as one step alongside RL Reinforcement learning on verifiable rewards: the reasoning stage
Preference optimisation needs a judge of better, and that judge — human or reward model — is where the biases come from. For some tasks no judge is needed. A maths problem has a checkable final answer; code either passes its tests or does not. Reasoning models are trained on exactly these tasks with reinforcement learning from verifiable rewards: sample an attempt, check it with a program, reward it if correct.
What the model learns from this is not new knowledge but a new habit. Because only the final answer is scored, the model is free to write as much working as it likes before answering, and attempts that check their own steps, back up after a mistake and try another route get rewarded more often. Longer, self-correcting chains of thought emerge because they win. DeepSeek reported this in its R1 work: RL on the base model alone, rewarded only for correct and well-formatted answers, produced long reasoning traces nobody had demonstrated.
The algorithm DeepSeek used, GRPO (group relative policy optimisation), also shows how simple the signal can be. For each problem, sample a group of attempts, reward each, and score every attempt against its own group:
Ai=std(r1,…,rG)ri−mean(r1,…,rG)Take a group of four attempts where one is correct, rewards [1,0,0,0]:
Textmean = 0.25 std = sqrt((0.75^2 + 3 x 0.25^2) / 4) = 0.433advantages: correct attempt (1 - 0.25) / 0.433 = +1.732 each wrong one (0 - 0.25) / 0.433 = -0.577The one correct attempt has every token's probability pushed up strongly; the three wrong ones are pushed down gently. The group mean replaces the separate value model PPO needs, which saves a model's worth of memory. And if all four attempts are right, or all wrong, every advantage is zero and the problem teaches nothing — so training data is chosen to be hard but not impossible for the current model.
The limits follow from the design. The reward only sees the final answer, so a lucky guess after muddled reasoning is rewarded, and a checker with a loophole (tests that can be special-cased, an answer format that can be gamed) will be found and exploited — reward hacking again, now against a program. And it only applies where answers can be checked, which is why reasoning models gained most in maths and code. Production pipelines combine all of these: SFT, then verifiable-reward RL, then preference optimisation for the parts no program can grade.
Adapting a model without retraining all of it
Full fine-tuning updates every weight, and the memory arithmetic makes clear why that is rarely feasible. For a 7B model with the Adam optimiser:
Textweights (bf16) 2 bytes/paramgradients (bf16) 2 bytes/paramAdam first moment (fp32) 4 bytes/paramAdam second moment (fp32) 4 bytes/paramfp32 master weights 4 bytes/param ---------------- 16 bytes/param6.74e9 x 16 = 108 GB - before a single activation is stored.LoRA avoids this by freezing the original weights and learning a low-rank correction alongside them. Instead of updating W directly, learn ΔW=BA where B is d×r and A is r×d, with r far smaller than d:
W′=W+rαBACount it for one 4096×4096 attention projection at rank 8:
TextFull weight matrix: 4096 x 4096 = 16,777,216 parametersLoRA A: 8 x 4096 = 32,768LoRA B: 4096 x 8 = 32,768 ------------LoRA total: 65,536 = 0.39% of the full matrixAcross Q, K, V, O in all 32 layers: 4 x 65,536 x 32 = 8,388,608 trainable parametersThat is 0.12% of the model's 6.74 billion.Memory: 13.5 GB frozen weights (2 bytes each, no optimiser state) + 0.13 GB for the LoRA parameters and their optimiser state = comfortably inside a single consumer accelerator.Why does a rank-8 correction suffice? Because the change needed to adapt a pretrained model to a new task is empirically low-rank, even though the model itself is not. The base weights already encode the capability; the adaptation is a small nudge in a low-dimensional subspace. Ranks of 8–64 cover the overwhelming majority of real fine-tuning tasks.
Two further practical wins: the adapters are tens of megabytes, so you can store hundreds of task-specific adapters against one shared base model and swap them at request time; and because the base weights are untouched, the original model is always recoverable.
Two failure modes to name
Catastrophic forgetting. Fine-tune aggressively on a narrow dataset and the model loses general ability. Train hard on legal contracts and general reasoning degrades — the weights that encoded it have been overwritten by whatever reduces loss on contracts. Defences: a low learning rate (1e-5 or below for full fine-tuning), few epochs, a mixture of general data blended into the fine-tuning set, or LoRA, which limits how far the weights can move by construction.
The alignment tax. Aligned models often score slightly worse on raw capability benchmarks than the base model they came from. Some of this is real: the KL penalty and the safety training constrain the output distribution. Some of it is measurement artefact — benchmarks that expect bare completions penalise a model that has learned to write full explanatory answers. Either way, expect the aligned model to be somewhat less capable on paper and considerably more useful in practice.
Choosing the right intervention
Almost every "should we fine-tune?" conversation is answered correctly by working down this list and stopping at the first row that works.
Approach Fixes Does not fix Cost Better prompting, few-shot examples Format, tone, task framing Missing knowledge; genuinely new capability Minutes Retrieval (RAG) Private, current or proprietary facts Behaviour, style, output format Days of engineering LoRA fine-tuning Consistent format, domain voice, structured output Facts that change often Hours of compute, hundreds of curated examples Full fine-tuning Deep domain shift, a very different language or modality — Days of compute, thousands of examples, forgetting risk Continued pretraining An entire domain the base model barely saw — Billions of tokens; only worth it at scale Pretraining from scratch Everything — Millions of dollars The most common expensive mistake is reaching for fine-tuning to inject facts. Fine-tuning teaches behaviour reliably and facts unreliably — a fact seen a handful of times during a short fine-tune is weakly encoded, easily contradicted by pretrained knowledge, and impossible to update without retraining. If the requirement is "the model should know our current product catalogue", that is a retrieval problem. If the requirement is "the model should always reply in our house style with these six sections", that is a fine-tuning problem. Getting that classification right saves most teams a month.