How Large Language Models Work

Fine-Tuning vs Pretraining vs RLHF


Download a raw pretrained model — one that has been trained on text and nothing else — and ask it a question:

Text
Prompt: What is the capital of France?Output: What is the largest country in Europe? What is the longest        river in the world? What is the highest mountain? Answers to        these and other questions can be found in our quiz pack...

The model is not broken. It is working perfectly. It was trained to continue text, and on the open internet, a line that looks like a quiz question is most often followed by more quiz questions. It gave the statistically best continuation. It just was not the one you wanted.

This gap — between a model that predicts text and a model that answers you — is what the training pipeline after pretraining exists to close. The classic pipeline has three stages, with three completely different data types, three cost profiles differing by four orders of magnitude. Confusing them is the single most common source of bad decisions about how to adapt a model, so it is worth being precise about what each one actually changes.

Three stages, and what each one actually addsPretraining: all of thecapability, trillions of tokensSFT: the shape of a reply, thousands of examplesPreference tuning: which valid reply to preferAdapters: task fitting without touching the base
The raw model completes text rather than answering, so what the later stages add is behaviour, not knowledge.

The three stages at a glance

PretrainingSupervised fine-tuningPreference optimisation
Question it answersHow does language work?What shape should a reply have?Which of two valid replies is better?
DataTrillions of tokens of raw text10k–1M prompt/response pairs10k–1M ranked comparisons
Labels come fromThe text itselfHumans writing or curating answersHumans choosing between outputs
Compute102110^{21}–102510^{25} FLOPs101710^{17}–101910^{19} FLOPs101810^{18}–102010^{20} FLOPs
Wall-clockWeeks to months on thousands of acceleratorsHours to days on 8Days, plus months of human labelling
What it addsKnowledge and capabilityFormat and instruction-followingJudgement about quality, tone, safety
Fraction of total compute~99%<1%<1%

Almost all of a model's ability is created in pretraining. Everything after it is teaching the model to use abilities it already has — which is why it can be done with a rounding error's worth of compute.

That last row describes the classic chat-model pipeline. Reasoning models added a fourth stage — reinforcement learning on problems whose answers can be checked, covered below — and it can absorb a much larger share of compute than the <1% shown here, because it is limited by how many problems you can generate and check, not by human labelling.

Pretraining: where the capability comes from

The objective is the plainest one available: predict the next token, and take the loss as −log⁡p-\log p of the true token. No labels, no annotation, no human in the loop. Any text is training data, which is what makes the scale possible.

The compute required follows a reliable approximation:

C≈6NDC \approx 6ND

with NN parameters and DD training tokens. The 6 comes from roughly 2 FLOPs per parameter per token in the forward pass and about 4 in the backward pass. Work a 7-billion-parameter model trained on 2 trillion tokens:

Text
C = 6 x 7e9 x 2e12 = 8.4e22 FLOPsAt an effective 1.2e14 FLOP/s per accelerator (about 40% of an A100's bf16 peak):  8.4e22 / 1.2e14 = 7.0e8 seconds = 8,100 accelerator-daysOn 1,024 accelerators: about 8 days of continuous running.

Add data collection, deduplication, filtering, failed runs and evaluation, and the realistic cost of a competent 7B model is measured in millions of dollars and a team of experienced people. This is the number that determines who gets to pretrain and who does not.

A crucial constraint shapes the design. Compute-optimal scaling work found that, for a fixed budget, model size and data should grow together at roughly 20 tokens per parameter. Earlier large models were badly undertrained on this measure — parameters were scaled aggressively while data lagged — and a smaller model trained on more tokens beat them. Models intended for heavy deployment go further still, training well past the compute-optimal point, because a smaller model costs less on every single inference call thereafter.

What exists after pretraining, and what does not

PresentAbsent
Grammar, syntax, discourse structureAny notion that a question should be answered
Broad world knowledgeA consistent persona or voice
Multiple languages and programming languagesRefusal of harmful requests
Arithmetic and code semanticsKnowing when to say "I don't know"
In-context learning from examples in the promptAny preference for helpful over merely plausible

The right-hand column is not a list of missing knowledge. It is a list of missing behaviours. That distinction is the key to everything that follows.

Supervised fine-tuning: teaching the shape of a reply

SFT continues the exact same next-token objective, on a different dataset: prompt–response pairs written or curated by people.

JSON
{  "messages": [    {"role": "user", "content": "Why does bread rise?"},    {"role": "assistant", "content": "Yeast in the dough consumes sugars and releases carbon dioxide. The gluten network formed when the dough is kneaded traps those gas bubbles, and the trapped gas expands the dough."}  ]}

One detail separates SFT from ordinary training, and getting it wrong quietly ruins the result: loss masking. The loss is computed only on the assistant's tokens. The user's tokens are set to the ignore index and contribute nothing.

Python
import torchIGNORE = -100def build_example(tokenizer, prompt, response):    p_ids = tokenizer(prompt,   add_special_tokens=False).input_ids    r_ids = tokenizer(response, add_special_tokens=False).input_ids    input_ids = p_ids + r_ids    labels    = [IGNORE] * len(p_ids) + r_ids   # only the response is learned    return torch.tensor(input_ids), torch.tensor(labels)

Skip the masking and you are training the model to generate user questions as well as answers — which is exactly the failure the whole stage exists to fix. Models fine-tuned without masking often start producing a plausible answer and then inventing the user's follow-up question.

Quality dominates quantity here

This is the finding that most surprises people coming from a pretraining mindset. Careful work has shown that a few thousand meticulously written, diverse examples can produce better instruction-following than hundreds of thousands of noisy, machine-generated ones. The reason follows from the table above: the model already possesses the capability. SFT is demonstrating a format, and a thousand clear demonstrations beat a hundred thousand muddled ones.

The practical implication is uncomfortable but useful: if your fine-tune is disappointing, adding more data of the same quality is usually the wrong response. Read fifty of your own examples carefully. The problem is normally in there.

Preference optimisation: choosing between valid answers

SFT teaches the model to produce an answer. It cannot teach it which of several correct answers is better, because the training signal only ever shows one target. Consider two responses to "my code throws a null pointer exception":

  • "A null pointer exception occurs when you dereference a null reference." — true, and useless.
  • "Something is null when you use it. Check the variable on the line the trace points to, and trace back to where it was last assigned — a function returning null on a failure path is the usual cause." — also true, and actually helpful.

Both are valid continuations. Nobody can easily write a loss function that prefers the second. But almost anyone can look at the pair and pick. That asymmetry — hard to specify, easy to judge — is what preference optimisation exploits.

The reward-model route (RLHF)

Three steps.

Collect comparisons. Sample several responses from the SFT model for each prompt, and have annotators rank them. The output is triples: prompt xx, preferred response ywy_w, rejected response yly_l.

Train a reward model. Take a copy of the model, replace the vocabulary-sized output head with a single scalar, and train it so preferred responses score higher. The loss comes from the Bradley–Terry model of pairwise choice:

LRM=−log⁡σ(rϕ(x,yw)−rϕ(x,yl))\mathcal{L}_{\text{RM}} = -\log \sigma\big(r_\phi(x, y_w) - r_\phi(x, y_l)\big)

Only the difference in scores appears, so the reward model never needs an absolute scale — it only needs to get orderings right.

Optimise the policy. Use reinforcement learning (classically PPO) to maximise reward, with a penalty that keeps the model from straying far from where it started:

objective=E[rϕ(x,y)]−β KL(πθ ∥ πref)\text{objective} = \mathbb{E}\big[r_\phi(x,y)\big] - \beta \,\mathrm{KL}\big(\pi_\theta \,\|\, \pi_{\text{ref}}\big)

The KL term is not optional. Without it the policy discovers that the reward model is an imperfect proxy and games it — a phenomenon called reward hacking. Typical symptoms are dryly funny: outputs balloon in length because annotators mildly preferred longer answers; every reply opens with effusive agreement because that scored well; the model produces confident-sounding structure with no content. The reward went up. The model got worse.

A reward model is a learned approximation of human judgement, and optimising hard against any approximation eventually optimises the gap rather than the goal.

The direct route (DPO)

The RLHF pipeline is heavy: three models in memory at once, an RL loop that is notoriously fiddly to stabilise. Direct Preference Optimisation showed that the reward model can be skipped entirely. The same optimum is reachable with a plain supervised loss on the preference pairs:

LDPO=−log⁡σ ⁣(βlog⁡πθ(yw∣x)πref(yw∣x)−βlog⁡πθ(yl∣x)πref(yl∣x))\mathcal{L}_{\text{DPO}} = -\log\sigma\!\left(\beta\log\frac{\pi_\theta(y_w\mid x)}{\pi_{\text{ref}}(y_w\mid x)} - \beta\log\frac{\pi_\theta(y_l\mid x)}{\pi_{\text{ref}}(y_l\mid x)}\right)

Read it as: raise the model's probability of the preferred response relative to the frozen reference, lower it for the rejected one, and let β\beta control how far from the reference the model may drift. No reward model, no sampling loop, no RL.

RLHF (PPO)DPO
Models needed in memoryPolicy, reference, reward, valuePolicy and reference
Training loopRL — sensitive to hyperparametersSupervised — stable
Can score new samples on the flyYes, via the reward modelNo — limited to the collected pairs
Typical useLarge training runs with continuous data collectionMuch open-model alignment work, often as one step alongside RL

Reinforcement learning on verifiable rewards: the reasoning stage

Preference optimisation needs a judge of better, and that judge — human or reward model — is where the biases come from. For some tasks no judge is needed. A maths problem has a checkable final answer; code either passes its tests or does not. Reasoning models are trained on exactly these tasks with reinforcement learning from verifiable rewards: sample an attempt, check it with a program, reward it if correct.

What the model learns from this is not new knowledge but a new habit. Because only the final answer is scored, the model is free to write as much working as it likes before answering, and attempts that check their own steps, back up after a mistake and try another route get rewarded more often. Longer, self-correcting chains of thought emerge because they win. DeepSeek reported this in its R1 work: RL on the base model alone, rewarded only for correct and well-formatted answers, produced long reasoning traces nobody had demonstrated.

The algorithm DeepSeek used, GRPO (group relative policy optimisation), also shows how simple the signal can be. For each problem, sample a group of attempts, reward each, and score every attempt against its own group:

Ai=ri−mean(r1,…,rG)std(r1,…,rG)A_i = \frac{r_i - \text{mean}(r_1,\dots,r_G)}{\text{std}(r_1,\dots,r_G)}

Take a group of four attempts where one is correct, rewards [1,0,0,0][1, 0, 0, 0]:

Text
mean = 0.25      std = sqrt((0.75^2 + 3 x 0.25^2) / 4) = 0.433advantages:  correct attempt   (1 - 0.25) / 0.433 = +1.732             each wrong one    (0 - 0.25) / 0.433 = -0.577

The one correct attempt has every token's probability pushed up strongly; the three wrong ones are pushed down gently. The group mean replaces the separate value model PPO needs, which saves a model's worth of memory. And if all four attempts are right, or all wrong, every advantage is zero and the problem teaches nothing — so training data is chosen to be hard but not impossible for the current model.

The limits follow from the design. The reward only sees the final answer, so a lucky guess after muddled reasoning is rewarded, and a checker with a loophole (tests that can be special-cased, an answer format that can be gamed) will be found and exploited — reward hacking again, now against a program. And it only applies where answers can be checked, which is why reasoning models gained most in maths and code. Production pipelines combine all of these: SFT, then verifiable-reward RL, then preference optimisation for the parts no program can grade.

Adapting a model without retraining all of it

Full fine-tuning updates every weight, and the memory arithmetic makes clear why that is rarely feasible. For a 7B model with the Adam optimiser:

Text
weights (bf16)              2 bytes/paramgradients (bf16)            2 bytes/paramAdam first moment (fp32)    4 bytes/paramAdam second moment (fp32)   4 bytes/paramfp32 master weights         4 bytes/param                           ----------------                           16 bytes/param6.74e9 x 16 = 108 GB - before a single activation is stored.

LoRA avoids this by freezing the original weights and learning a low-rank correction alongside them. Instead of updating WW directly, learn ΔW=BA\Delta W = BA where BB is d×rd \times r and AA is r×dr \times d, with rr far smaller than dd:

W′=W+αrBAW' = W + \frac{\alpha}{r}BA

Count it for one 4096×40964096 \times 4096 attention projection at rank 8:

Text
Full weight matrix:   4096 x 4096 = 16,777,216 parametersLoRA A:               8 x 4096    =     32,768LoRA B:               4096 x 8    =     32,768                                   ------------LoRA total:                             65,536   = 0.39% of the full matrixAcross Q, K, V, O in all 32 layers:  4 x 65,536 x 32 = 8,388,608 trainable parametersThat is 0.12% of the model's 6.74 billion.Memory: 13.5 GB frozen weights (2 bytes each, no optimiser state)      +  0.13 GB for the LoRA parameters and their optimiser state      = comfortably inside a single consumer accelerator.

Why does a rank-8 correction suffice? Because the change needed to adapt a pretrained model to a new task is empirically low-rank, even though the model itself is not. The base weights already encode the capability; the adaptation is a small nudge in a low-dimensional subspace. Ranks of 8–64 cover the overwhelming majority of real fine-tuning tasks.

Two further practical wins: the adapters are tens of megabytes, so you can store hundreds of task-specific adapters against one shared base model and swap them at request time; and because the base weights are untouched, the original model is always recoverable.

Two failure modes to name

Catastrophic forgetting. Fine-tune aggressively on a narrow dataset and the model loses general ability. Train hard on legal contracts and general reasoning degrades — the weights that encoded it have been overwritten by whatever reduces loss on contracts. Defences: a low learning rate (1e-5 or below for full fine-tuning), few epochs, a mixture of general data blended into the fine-tuning set, or LoRA, which limits how far the weights can move by construction.

The alignment tax. Aligned models often score slightly worse on raw capability benchmarks than the base model they came from. Some of this is real: the KL penalty and the safety training constrain the output distribution. Some of it is measurement artefact — benchmarks that expect bare completions penalise a model that has learned to write full explanatory answers. Either way, expect the aligned model to be somewhat less capable on paper and considerably more useful in practice.

Choosing the right intervention

Almost every "should we fine-tune?" conversation is answered correctly by working down this list and stopping at the first row that works.

ApproachFixesDoes not fixCost
Better prompting, few-shot examplesFormat, tone, task framingMissing knowledge; genuinely new capabilityMinutes
Retrieval (RAG)Private, current or proprietary factsBehaviour, style, output formatDays of engineering
LoRA fine-tuningConsistent format, domain voice, structured outputFacts that change oftenHours of compute, hundreds of curated examples
Full fine-tuningDeep domain shift, a very different language or modality—Days of compute, thousands of examples, forgetting risk
Continued pretrainingAn entire domain the base model barely saw—Billions of tokens; only worth it at scale
Pretraining from scratchEverything—Millions of dollars

The most common expensive mistake is reaching for fine-tuning to inject facts. Fine-tuning teaches behaviour reliably and facts unreliably — a fact seen a handful of times during a short fine-tune is weakly encoded, easily contradicted by pretrained knowledge, and impossible to update without retraining. If the requirement is "the model should know our current product catalogue", that is a retrieval problem. If the requirement is "the model should always reply in our house style with these six sections", that is a fine-tuning problem. Getting that classification right saves most teams a month.