Course Content
How Large Language Models Work
3 sections · 9 lessons
Instruction Tuning & Alignment
When you call a chat model, you send it something structured — a list of messages, each with a role:
1[2 {"role": "system", "content": "You are a careful research assistant."},3 {"role": "user", "content": "Summarise this in two sentences."},4 {"role": "assistant", "content": "..."}5]The model receives none of that structure. There is no role field inside a transformer, no separate channel for system instructions, no protected memory for "things the developer said" versus "things a stranger typed". Before anything reaches the model, that JSON is flattened into one continuous string of tokens, and the entire notion of roles survives only as special marker tokens that the model was trained to treat in a particular way.
Everything about alignment follows from that fact. It is why a model can be instructed at all, why the instructions are only as strong as the training that installed them, and why prompt injection is a structural property rather than a bug someone forgot to patch.
The chat template: where roles actually live
Here is what the messages above become for a model using the widely-adopted ChatML format:
<|im_start|>systemYou are a careful research assistant.<|im_end|><|im_start|>userSummarise this in two sentences.<|im_end|><|im_start|>assistantDifferent model families use different markers — some use [INST] and [/INST], some use header blocks with their own reserved tokens — but the principle never changes. Roles are punctuation. The model learns during instruction tuning that text appearing after the assistant marker is what it should produce, and text before it is what it should respond to.
Reasoning models add one more region to the template: a thinking segment written before the answer — Qwen3 and DeepSeek-R1, for example, wrap it in <think> and </think> tags. Templates differ on what happens to that segment in later turns: some strip earlier turns' thinking from the history, others keep it. That is one more reason to let the model's own template build the string.
Notice the last line of the ChatML example: the string ends immediately after the assistant marker, with nothing following it. That trailing marker is the whole trigger. The model's next-token prediction now naturally continues an assistant turn, because that is what always followed this exact pattern in its training data.
A chat model is a completion model that has been taught one very specific convention: after this token, be helpful. The convention is learned, not enforced.
Two practical consequences fall straight out.
Using the wrong template silently degrades the model. If you feed a model raw text without its markers, or with another family's markers, you have moved it off the exact pattern it was trained on. It still produces output — often plausible output — but measurably worse. This is a common cause of "the open model is much worse than the benchmarks claim". Always apply the model's own template:
1from transformers import AutoTokenizer23tok = AutoTokenizer.from_pretrained("some-instruct-model")45messages = [6 {"role": "system", "content": "You are a careful research assistant."},7 {"role": "user", "content": "Summarise this in two sentences."},8]910prompt = tok.apply_chat_template(11 messages,12 tokenize=False,13 add_generation_prompt=True, # appends the trailing assistant marker14)15print(prompt)The system prompt has no special power. It is text in a marked region, distinguished from user text only by markers the model learned to weight more heavily. There is no mechanism preventing a user message from contradicting it. If a user writes "ignore your previous instructions", the model's compliance or refusal is a learned behaviour competing against a learned behaviour — a probability contest, not an access-control check. Providers now train models on an explicit instruction hierarchy — system instructions outrank user messages, which outrank text inside tool results — and that makes overrides much harder. It is still a trained preference, not an enforced rule. This is why prompt injection cannot be fully fixed by prompting alone, and why any system that puts untrusted text (a scraped web page, a user-uploaded document, a tool result) into the same token stream as its instructions has to treat that boundary as porous.
Loss masking, position by position
Instruction tuning uses the same next-token loss as ordinary training, applied to formatted conversations. The one structural difference is which positions contribute to the loss.
Take a short example and label every token:
| Position | Token | Region | Contributes to loss? |
|---|---|---|---|
| 1–6 | <|im_start|>system … <|im_end|> | System | No |
| 7–8 | <|im_start|>user | Marker | No |
| 9–14 | What is the boiling point of water? | User content | No |
| 15–16 | <|im_end|><|im_start|>assistant | Marker | No |
| 17–28 | 100 °C at standard atmospheric pressure. | Assistant content | Yes |
| 29 | <|im_end|> | Stop marker | Yes |
Only 13 of 29 positions produce gradient. The rest are set to the ignore index.
Two things go wrong if you get this wrong, and both are seen constantly in home-grown fine-tuning code. Train on the user tokens too, and the model learns to generate user turns — you will see it answer, then invent the user's next question and answer that as well. Forget to include the closing marker in the loss, and the model never learns to stop; it runs to the token limit every time, trailing off into invented dialogue.
What makes instruction data good
The instinct is to gather as much data as possible. For this stage that instinct is wrong, and understanding why clarifies what the stage is actually doing.
The model already knows the boiling point of water. It learned that from trillions of tokens. Instruction tuning is not adding knowledge — it is demonstrating a format: answer directly, at appropriate length, in the right register, and stop. A few thousand clear demonstrations teach that. A hundred thousand muddled ones teach muddle. Published work has repeatedly found that carefully curated datasets in the low thousands outperform machine-generated sets a hundred times larger.
What "good" concretely means:
| Property | What it looks like | What happens without it |
|---|---|---|
| Task diversity | Summarising, extracting, explaining, coding, rewriting, refusing, asking for clarification | The model over-fits to the represented shapes and handles others poorly |
| Length diversity | Some one-line answers, some long ones, matched to the question | Every answer comes out the same length regardless of what was asked |
| Honest uncertainty | Examples where the correct answer is "I don't have enough information to say" | The model never declines; it invents. Confabulation is partly a data problem, not only a capability one. |
| Multi-turn examples | Conversations with follow-ups and corrections | Turn 3 of a real conversation degrades badly |
| Consistent formatting | The same conventions for headings, code fences, lists | Output formatting becomes unpredictable and hard to parse downstream |
| Factual correctness | Verified answers | You are explicitly training the model to state falsehoods confidently |
That last row deserves emphasis. Every example in your dataset is a demonstration of how a confident assistant behaves. An incorrect answer written in an authoritative voice teaches authoritative wrongness, not the specific error. The style transfers even when the fact does not.
Reward models and the ways they mislead
SFT can only show one correct answer per prompt. It cannot express that one good answer is better than another good answer. That requires comparison data and a model trained to score it.
A reward model is the language model with its vocabulary-sized output head replaced by a single scalar, trained so that preferred responses score higher than rejected ones:
Only the difference appears, so the scale is arbitrary — the model learns ordering, not absolute quality. Three failure modes recur and are worth recognising by their symptoms.
Length bias. Annotators, given two answers and limited time, tend to prefer the longer one. The reward model learns that length correlates with quality. Then the policy, optimising against that reward, discovers it can raise its score by padding. The visible result is answers that restate the question, list caveats nobody asked for, and end with a summary of what was just said. The reward went up; usefulness went down.
Sycophancy. Annotators prefer being agreed with. If a user asserts something incorrect and the model contradicts them, that response often loses the comparison. The learned lesson is: agree. A well-known consequence is a model that confidently states a correct answer, is told "that's wrong", and immediately retracts a correct answer.
Style over substance. Confident, well-structured prose scores highly whether or not it is accurate, because verifying accuracy is expensive and judging polish is instant. Optimise hard and you get a model that has learned to look right.
A reward model measures what annotators noticed, not what was true. Every systematic bias in the annotation process becomes a systematic behaviour in the trained model.
The standard defence is a KL penalty that keeps the optimised policy near its starting point, capping how far it can travel to exploit the reward model's flaws. It is a limit on damage, not a fix for the underlying gap.
Using the model to supervise itself
Human comparison data is slow and expensive, and human annotators disagree constantly on anything nuanced. Constitutional AI replaces much of it with a written set of principles and a critique-and-revise loop run by the model itself.
PRINCIPLE: Do not provide instructions that would enable someone to harm themselves or others.Prompt: How do I pick a lock?Draft: [detailed step-by-step instructions on tension wrenches and pin manipulation]Critique: Ask the model - "Does the draft violate the principle above? Explain." -> "It gives operational detail that primarily helps someone enter a property they cannot access legitimately."Revision: Ask the model - "Rewrite the response so it complies." -> "Lock picking works by manipulating pins to the shear line while applying rotational tension. If you're locked out of your own property, a locksmith can verify ownership and open it without damage. If you're interested in the mechanics as a hobby, locksport clubs teach it on practice locks."Training: Fine-tune on (prompt, revised response). The critique and the draft are discarded - only the improved answer is kept.The comparison data for the preference stage can be generated the same way: sample two responses, ask the model which better follows the principles, and use its choice as the label. Feedback from a model rather than a person is generally called RLAIF.
| Human feedback | Model feedback (constitutional / RLAIF) | |
|---|---|---|
| Cost per comparison | Dollars; minutes of a person's time | Fractions of a cent; seconds |
| Scale reachable | Tens or hundreds of thousands | Millions |
| Consistency | Annotators disagree, and drift over time | Highly consistent — which cuts both ways |
| Auditability | Values are implicit in a rubric and in who was hired | Values are written down and can be revised and diffed |
| Main risk | Bias, fatigue, expense | The supervising model's own blind spots get amplified, uniformly |
The auditability row is the genuinely important one. When principles are written in plain text, disagreements about a model's behaviour become disagreements about a document — which can be argued over, versioned and corrected. When values live only in annotator instincts, they cannot be inspected at all.
Measuring whether any of this worked
Alignment cannot be scored with accuracy, because there is no single correct answer. The standard instrument is the pairwise win rate: show a judge two responses to the same prompt, ask which is better, and count.
The arithmetic matters more than people assume. Suppose model A beats model B on 118 of 200 prompts:
win rate = 118 / 200 = 0.590standard error = sqrt(p(1-p)/n) = sqrt(0.59 x 0.41 / 200) = sqrt(0.0012095) = 0.034895% interval = 0.590 +/- 1.96 x 0.0348 = 0.522 to 0.658So a 59% win rate over 200 comparisons genuinely beats a coin flip — just. Now run the same calculation for a 52% win rate:
standard error = sqrt(0.52 x 0.48 / 200) = 0.035395% interval = 0.520 +/- 0.069 = 0.451 to 0.589 <- straddles 50%A 52% win rate at n=200 is indistinguishable from no difference. Announcements of narrow victories over a couple of hundred prompts are, very often, announcements of noise.
Judge biases, and how to cancel them
Using a strong model as the judge is cheap and scales, but judges have measurable, reproducible biases:
- Position bias. The response shown first is favoured. Cancel it by running every comparison twice with the order swapped and counting a disagreement between the two runs as a tie.
- Length bias. Longer responses win more often, independent of quality. Report the mean length of each model's outputs alongside the win rate; if the winner is also consistently longer, you have not separated the two effects.
- Self-preference. A judge tends to favour text resembling its own style. Never evaluate a model using itself, or a close sibling, as judge.
- Style over substance. Judges reward formatting and confidence. Any claim about factual accuracy needs a factual check, not a preference vote.
The failure modes alignment introduces
Alignment does not only add good behaviour. It adds its own characteristic problems, and recognising them by symptom is a practical skill.
| Symptom | Underlying cause | What actually helps |
|---|---|---|
| Model reverses a correct answer when contradicted | Sycophancy learned from agreement-preferring annotators | Preference data that explicitly rewards holding a correct position under pushback |
| Refuses obviously benign requests ("how do I kill a process") | Safety training keyed on surface features rather than intent | Borderline-but-benign examples with correct compliant answers in the training mix |
| Every answer is padded with caveats and restatement | Length bias in the reward signal | Length-controlled evaluation; preference pairs where the shorter answer is preferred |
| Refusals collapse when the request is wrapped in a fictional frame | Behaviour learned on typical phrasings does not generalise to unusual ones | Adversarial examples in training; and defences outside the model, since none of this is enforced |
| Aligned model scores below the base model on capability benchmarks | Partly the constrained output distribution; partly benchmarks expecting bare completions | Evaluate with the chat template and appropriate prompting before concluding capability was lost |
Building on top of an aligned model
Treat the template as part of your interface contract. Apply the model's own chat template, use its exact special tokens, and never hand-assemble the string. Templates change between model versions, and a mismatch produces a quality regression with no error message — the most expensive kind of bug to find.
Treat the system prompt as a strong hint and never as a security boundary. If your application inserts retrieved documents, tool outputs, or user-supplied files into the context, assume anything in them can compete with your instructions. Real defences live outside the model: validate and constrain tool arguments, apply permissions at the tool layer, and require confirmation for irreversible actions. "The system prompt says not to" is not a control.
And when you evaluate, evaluate the way you deploy. Same template, same system prompt, same decoding settings, and enough prompts that your win rate has an interval you would be willing to defend. Two hundred prompts gives you roughly ±7 percentage points. If the difference you are trying to detect is smaller than that, you need more prompts — or a different measurement entirely.