Course Content
Introduction to Generative AI
3 sections · 9 lessons
Transformers as Generative Models
Take the sentence: "The trophy would not fit in the brown suitcase because it was too large." What does it refer to?
The trophy, obviously. Change one word — "because it was too small" — and now it is the suitcase. A single adjective eleven words away flips the answer.
Now build a network that gets this right. The natural approach for years was a recurrent network: read one word at a time, keep a hidden state vector, update it at each word. By the time you reach it, everything the model knows about the preceding eleven words has to be inside one fixed-size vector — say 512 numbers.
That vector has been overwritten eleven times. Each update mixes in new information and, necessarily, degrades what was already there. "Trophy" entered at position 2 and has survived nine rounds of being blended with other content. For a 12-word sentence this is merely difficult. For a 2,000-word document it is hopeless: the model is trying to compress an entire novel's worth of context into a vector the size of a small image thumbnail.
There is a second problem, and commercially it was the more serious one. Step t of a recurrent network cannot be computed until step t−1 finishes. A 1,000-token training sequence requires 1,000 sequential operations during training, when the correct answer is already known for every position. A machine with thousands of parallel cores sits mostly idle.
Both problems have the same root: the model is compressing the past. So stop compressing it. Keep every previous position available, and at each position look up whatever is relevant. That is attention, and it dissolves both problems at once.
Recurrence forces information through a fixed-size bottleneck at every step. Attention removes the bottleneck by keeping the past addressable instead of summarised.
Attention as a soft lookup
Think of a dictionary. You have a query, entries have keys, and each key maps to a value. A normal lookup matches one key exactly. Attention does a soft version: compare the query against every key, convert the similarities into weights that sum to one, and return the weighted average of all values.
Every token produces all three vectors from its own representation, through learned matrices:
Loosely: the query is "what am I looking for", the key is "what I can offer", and the value is "what I actually contribute if selected". Then:
Working through one position
Take the token it in the sentence above. Its query vector is compared against the keys of every earlier token. Suppose the raw dot products come out as:
| Token | q⋅k | Scaled by dk=8 | exp(⋅) | Attention weight |
|---|---|---|---|---|
| trophy | 24.0 | 3.00 | 20.09 | 0.855 |
| suitcase | 6.0 | 0.75 | 2.12 | 0.090 |
| large | 2.0 | 0.25 | 1.28 | 0.055 |
The output at this position is 0.855×vtrophy+0.090×vsuitcase+0.055×vlarge. The representation of it now carries mostly the trophy's content. Note that no distance appears anywhere in that computation — the trophy being eleven positions back cost exactly nothing. Attention is distance-free, which is precisely what recurrence could not offer.
Distance-free has a side effect: attention on its own cannot tell word order. "Dog bites man" and "man bites dog" would look identical. So position is added back in. The original transformer added a fixed sinusoidal vector for each position to the token embeddings; most current language models instead use rotary position embeddings (RoPE), which rotate the query and key vectors by an angle that depends on position, so the dot product carries how far apart two tokens are.
Why divide by dk
This looks like a fussy detail and is not. If q and k have dk components each with roughly unit variance, their dot product has variance dk, so with dk=64 the typical magnitude is around 8, and extremes reach 20 or 30.
Feed differences of that size into a softmax and it saturates: one weight becomes essentially 1 and the rest essentially 0. A saturated softmax has near-zero gradient everywhere, so the model stops learning which positions to attend to. Dividing by dk brings the scores back to a range where the softmax is soft and its gradients are alive. Remove that one factor and large transformers do not train.
Multiple heads
A single attention operation produces one weighted average — one relationship per position. But it needs to resolve a pronoun, and simultaneously the model needs to track that fit is the main verb, that brown modifies suitcase, and so on.
So run several attention operations in parallel with separate learned projections, each in a smaller subspace, and concatenate the results. With a model dimension of 512 and 8 heads, each head works in 64 dimensions. Total compute is essentially unchanged; what you gain is eight different relationships computed at once. Trained models reliably develop heads that specialise — some tracking syntactic dependencies, some tracking the previous occurrence of the current token, some doing something no one has yet named.
Causal masking
A generative language model predicts the next token from the previous ones. If attention at position 5 could see position 6, the training objective would be trivially solvable by copying — the model would score perfectly during training and produce nothing at generation time, when the future does not exist.
The fix is a mask applied to the score matrix before the softmax. Every entry where the key position is later than the query position is set to −∞, which becomes exactly zero after exponentiation.
1import torch23def causal_attention(q, k, v):4 d_k = q.size(-1)5 scores = q @ k.transpose(-2, -1) / d_k ** 0.5 # (B, heads, T, T)6 T = scores.size(-1)7 mask = torch.triu(torch.ones(T, T, device=q.device), diagonal=1).bool()8 scores = scores.masked_fill(mask, float('-inf')) # block the future9 return torch.softmax(scores, dim=-1) @ vThis is what makes training efficient in a way recurrence never could be. With the mask in place, every position's prediction is computed in the same forward pass, and every position is trained on simultaneously. A 1,000-token sequence produces 1,000 training signals from one pass. The recurrent network needed 1,000 sequential steps to produce the same thing.
The transformer's decisive advantage is not that it attends. It is that masking lets it train on every position in parallel — turning a sequential problem into a matrix multiplication.
The full block, and the two architectures
One transformer block is attention plus a position-wise feed-forward network, each wrapped in a residual connection and a normalisation layer:
1class Block(nn.Module):2 def __init__(self, d, heads):3 super().__init__()4 self.ln1 = nn.LayerNorm(d)5 self.attn = MultiHeadAttention(d, heads)6 self.ln2 = nn.LayerNorm(d)7 self.mlp = nn.Sequential(8 nn.Linear(d, 4 * d), nn.GELU(), nn.Linear(4 * d, d)9 )1011 def forward(self, x):12 x = x + self.attn(self.ln1(x)) # mix across positions13 x = x + self.mlp(self.ln2(x)) # process each position14 return xThe division of labour is clean. Attention moves information between positions; the feed-forward network transforms each position independently. Stacking these alternates communication and computation. The normalisation goes before each sublayer rather than after — a small change from the original design that makes deep stacks train far more reliably, because it leaves the residual path clean.
Current open models such as Llama change a few parts: RMSNorm instead of LayerNorm, a gated feed-forward layer (SwiGLU) instead of the plain GELU one, and rotary positions. Many of the largest models also use mixture-of-experts layers, where each token is sent to only a few of many feed-forward blocks, so the parameter count grows much faster than the compute per token. The shape of the block, attention then feed-forward with residuals, is unchanged.
Two arrangements dominate.
| Decoder-only | Encoder–decoder | |
|---|---|---|
| Structure | One stack, causally masked throughout | Bidirectional encoder plus causal decoder with cross-attention |
| Input handling | Prompt and output are one sequence | Input encoded separately, attended to by the decoder |
| Best at | Open-ended generation, dialogue, code, few-shot prompting | Fixed input-to-output transforms: translation, summarisation |
| Advantage | Simpler; every token is a training signal; scales cleanly | Encoder sees the input bidirectionally, which helps when the input is fixed |
Decoder-only won for general-purpose models, and the reason is training efficiency rather than expressive power. In an encoder–decoder, only the decoder's tokens produce next-token predictions; in a decoder-only model, every token in the sequence does. Given a fixed compute budget that difference compounds enormously.
Training
The objective is one line: maximise the probability of each actual next token.
1logits = model(tokens[:, :-1]) # predict position t from t-1 and earlier2loss = F.cross_entropy(3 logits.reshape(-1, vocab_size),4 tokens[:, 1:].reshape(-1) # targets are the inputs, shifted by one5)No labels, no annotators. The text supplies its own targets, which is why this scaled to trillions of tokens while supervised datasets stalled in the millions.
The loss is usually reported as perplexity, exp(L), which has an interpretation worth carrying: it is roughly the number of options the model is effectively choosing between. Perplexity 1 means certainty; perplexity 50,000 on a 50,000-token vocabulary means the model has learned nothing. Real language models sit in the range of about 3 to 20 depending on the text — meaning that at a typical position the model has narrowed 50,000 possibilities down to a handful.
Scaling
The empirical finding that reshaped the field is that loss falls as a smooth power law in model size, data size and compute — straight lines on log-log axes, holding over many orders of magnitude. No plateau appeared where everyone expected one.
The practical form of the result concerns how to split a fixed compute budget. Early large models were heavily over-parameterised relative to their training data. The corrected guidance is roughly 20 training tokens per parameter for compute-optimal training.
| Parameters | Compute-optimal tokens | Note |
|---|---|---|
| 1B | ~20B | Trainable on a small cluster |
| 7B | ~140B | Common open-model scale |
| 70B | ~1.4T | Serious infrastructure required |
One qualification matters in practice: compute-optimal is about training cost, not deployment cost. A model that will serve billions of requests is worth training well past the optimal point, because a smaller model trained on more data is cheaper for every inference thereafter. Most widely deployed models are deliberately "over-trained" for this reason. Meta, for example, trained the 8-billion-parameter Llama 3 on more than 15 trillion tokens, close to a hundred times the 160 billion the table suggests.
The other observation from scaling is emergence: some capabilities are absent at small scale, then appear over a narrow range of size. Multi-step arithmetic, following instructions given only as examples in the prompt, and chain-of-thought reasoning all behave this way. There is legitimate debate about how much of this is a real phase change and how much is an artefact of scoring with all-or-nothing metrics — a model that is gradually getting better at a 5-step calculation scores zero on every step until it gets all five right. Either way, the practical lesson holds: a capability absent from a small model is not evidence that the approach cannot do it.
Since 2024 a second axis has mattered as much as model size: compute spent at answer time. Reasoning models are trained, largely with reinforcement learning on problems whose answers can be checked, to generate a long chain of intermediate reasoning before the final answer. Letting them think for longer tends to improve results on maths, coding and planning, at the price of more tokens, more latency and a bigger bill. Mechanically nothing new is happening: it is still next-token prediction, spending extra tokens on the working.
Decoding: how the text actually comes out
The model outputs a probability distribution over the whole vocabulary. Converting that into a token is a separate decision, and it changes output character more than most people expect.
| Strategy | Rule | Character | Use for |
|---|---|---|---|
| Greedy | Always take the highest-probability token | Deterministic; falls into repetition loops | Short factual answers |
| Beam search | Keep k candidate sequences, pick the most likely overall | Higher total probability, oddly bland text | Translation, not open generation |
| Temperature | Divide logits by T before softmax | T<1 sharpens, T>1 flattens | The main creativity dial |
| Top-k | Sample from the k most likely tokens | Blocks the long tail of nonsense | General use; k around 40 |
| Top-p (nucleus) | Sample from the smallest set whose probability sums to p | Adapts: narrow when the model is confident, wide when it is not | The usual default; p around 0.9 |
Hosted APIs usually expose only some of these, typically temperature and top-p, and some reasoning models accept only the defaults or reject the parameters entirely. The strategies still describe what happens inside; you just may not hold the dials.
Two of these are worth understanding rather than just setting.
Why beam search produces bland text. It finds high-probability sequences, and the highest-probability continuation of almost anything is generic — safe, common phrasing. Human writing is not maximally probable; it contains surprise. Beam search systematically removes exactly that. It works for translation, where there really is one best answer, and fails for open-ended writing.
Why top-p beats top-k. A fixed k is wrong in both directions. After "The capital of France is", the model puts perhaps 0.98 on "Paris" — and top-40 sampling still keeps 39 wrong alternatives available. After "She opened the door and saw", hundreds of continuations are reasonable, and k=40 arbitrarily discards most of them. Top-p sizes the candidate set to the model's own confidence: it might keep one token in the first case and three hundred in the second.
The cost of generation
Generation is sequential — token n+1 requires token n — which is a real limitation. Naively, generating 1,000 tokens means 1,000 forward passes over an ever-growing sequence, at O(n2) attention cost each time.
The KV cache is what makes this tractable. Since keys and values for earlier positions never change once computed, cache them and compute only the new position's query against the cached keys. Cost per token drops from quadratic to linear.
The trade is memory. Cache size is roughly 2×layers×heads×head dim×sequence length×bytes, per sequence. For a large model at long context this reaches many gigabytes — which is why serving long-context models is memory-bound rather than compute-bound, and why techniques that shrink the cache by sharing keys and values across heads are now standard.
Why one architecture ended up everywhere
| Transformer | Recurrent network | Diffusion | GAN | |
|---|---|---|---|---|
| Training parallelism | Full | None across time | Full | Full |
| Long-range dependencies | Direct, distance-free | Degrades with distance | Via attention blocks | Limited |
| Passes per sample | One per token | One per token | 10–1000 for the whole output | 1 |
| Exact likelihood | Yes | Yes | Bound only | No |
| Natural data type | Discrete sequences | Discrete sequences | Continuous data | Continuous data |
| Scaling behaviour | Smooth and predictable | Saturates | Good | Difficult |
The row that decided everything is the first. An architecture that saturates a GPU cluster during training gets to consume more data and more parameters per unit of wall-clock time, and the scaling laws convert that directly into capability. Attention's quality advantage on long dependencies is real, but parallel training is what let the field find out how far the approach goes.
Its reach also extends past text. Once data is cut into a sequence of tokens, the same machinery applies — image patches, audio codec codes, protein residues, actions in a control policy. The tokeniser changes; the model does not. This is how multimodal assistants work: an image is cut into patches, each patch becomes a vector, and those vectors sit in the same sequence as the words of your question.
What this means when you build something
Context is memory, and it is finite and quadratic. The model knows only what is in its context window plus whatever its weights encode. Doubling the prompt roughly quadruples attention cost and linearly increases cache memory. Retrieving the three relevant paragraphs beats pasting the whole manual, on cost and on quality — a long prompt full of irrelevant text dilutes the attention that should be landing on what matters.
Match decoding parameters to the task, not to a house default. Extraction, classification and code want low temperature and tight top-p where the model allows it: there is a right answer and variation is pure risk. For extraction, a structured-output mode that enforces a JSON schema does more than any temperature setting. Brainstorming and drafting want higher temperature: variation is the product. Setting one global temperature across an application with both kinds of call is a common and avoidable mistake.
Early tokens are load-bearing. Generation conditions on its own output, so a mistake at token 20 becomes an established premise for tokens 21 onwards, and the model will build on it coherently and confidently. This is why asking for reasoning before a conclusion works — the intermediate steps become context the final answer must be consistent with. It is also why a wrong first sentence rarely recovers, and why regenerating usually beats trying to correct mid-stream.
Fluency is not knowledge. The training objective rewards likely text. A confident, well-formed, entirely invented citation is exactly what a model optimised for likelihood should produce when it has no relevant knowledge, because plausible-looking citations are what its training data contains. Any application that depends on facts needs those facts supplied in the context and verified on the way out — the model's tone will be identical either way.