Course Content
Attention Mechanisms and Transformers
4 sections · 11 lessons
Positional Encoding & Residual Connections
Here is an experiment you can run in five minutes. Build a self-attention layer, feed it the token embeddings for dog bites man, and record the output. Now shuffle the input rows to man bites dog and run it again. Compare.
The outputs are identical — just permuted the same way the input was. Not similar: bit-for-bit the same vectors, in a different order. The layer has no idea the order changed, because nothing in softmax(QK⊤/dk)V references a position index. Scores depend only on which vectors are present, so a permutation of the input produces exactly the corresponding permutation of the output.
That property has a name — permutation equivariance — and it makes raw self-attention a very expensive bag of words. A model built this way will get word identity right and word order consistently wrong, and no amount of extra layers will fix it, because stacking permutation-equivariant layers gives you another permutation-equivariant function.
There are two separate structural problems to solve before a stack of attention layers becomes a working transformer. The first is telling the model where each token is. The second is making a deep stack of these layers trainable at all. This lesson does both, with the arithmetic.
Part 1: injecting position
Why it must be added to the input
Since attention cannot see position, position must be baked into the vectors it does see. The standard approach adds a position-dependent vector to each token embedding before the first layer:
Adding rather than concatenating is a deliberate choice. Concatenation would preserve the two signals in separate coordinates but widen every matrix in the model. Addition keeps dmodel fixed and relies on the model learning to separate the two contributions — which it can, because in high dimensions two random subspaces are nearly orthogonal, and because the projections WQ and WK are free to attend to whichever coordinates carry which signal.
Option A: learned absolute embeddings
The simplest thing that works: a lookup table with one trainable vector per position.
1class LearnedPositionalEmbedding(nn.Module):2 def __init__(self, max_len, d_model):3 super().__init__()4 self.pos_emb = nn.Embedding(max_len, d_model)56 def forward(self, x): # x: (B, T, d_model)7 T = x.size(1)8 positions = torch.arange(T, device=x.device)9 return x + self.pos_emb(positions) # broadcasts over the batchParameter cost: max_len × d_model. For BERT-base that is 512×768=393,216 parameters — about 0.36% of its 110 M total, which is negligible.
The hard limit is in the shape of that table. Position 512 has a row; position 513 does not exist. Feed a longer sequence and you get an index error, or worse, a silent wrap-around if someone wrote a modulo. There is no principled way to extrapolate, because position 513's embedding was never trained and there is no formula connecting it to position 512's.
Option B: sinusoidal encoding
The original transformer used a fixed formula instead. For position pos and dimension index i:
Dimensions are paired: each pair (2i,2i+1) is a sine and cosine at one frequency, and the frequency decreases geometrically as i increases. Pair 0 oscillates with wavelength 2π≈6.3 positions; the last pair has wavelength 2π⋅10000≈62,832 positions. Together they act like a clock face with hands of many different speeds — the fast hands distinguish neighbours, the slow hands distinguish regions of the sequence.
Compute it by hand for dmodel=4. There are two pairs. For pair i=0 the divisor is 100000=1; for pair i=1 it is 100002/4=100000.5=100.
| pos | dim 0: sin(pos) | dim 1: cos(pos) | dim 2: sin(pos/100) | dim 3: cos(pos/100) |
|---|---|---|---|---|
| 0 | 0.0000 | 1.0000 | 0.0000 | 1.00000 |
| 1 | 0.8415 | 0.5403 | 0.0100 | 0.99995 |
| 2 | 0.9093 | −0.4161 | 0.0200 | 0.99980 |
| 3 | 0.1411 | −0.9900 | 0.0300 | 0.99955 |
Every row is different, so positions are distinguishable. But the interesting property is what happens when you take dot products between rows.
The same value. Positions 1-and-2 and positions 2-and-3 have identical similarity, because both pairs are one step apart. This is not a coincidence — it falls out of the angle-difference identity:
so the dot product over all pairs is ∑icos(100002i/dΔ), a function of the offset Δ alone. For Δ=1 here that is cos(1)+cos(0.01)=0.5403+0.99995=1.54025, matching both computations above.
There is a stronger property. PE(pos+k) is a fixed linear function of PE(pos) — the same function for every pos. Within one frequency pair, write the rotation matrix
Then
Because WQ and WK are linear maps, the model can in principle learn to implement "attend to whatever is 3 positions back" as a single linear operation that works at every position. That is what makes sinusoidal encoding more than an arbitrary set of distinct labels.
Sinusoidal encoding gives every pair of positions a similarity that depends only on their distance, and makes relative offsets expressible as fixed linear maps — which is why a formula beats a lookup table despite having zero parameters.
1import math2import torch3import torch.nn as nn45class SinusoidalPositionalEncoding(nn.Module):6 def __init__(self, d_model, max_len=5000, dropout=0.1):7 super().__init__()8 self.dropout = nn.Dropout(dropout)910 pe = torch.zeros(max_len, d_model)11 position = torch.arange(max_len).unsqueeze(1).float() # (max_len, 1)1213 # 1 / 10000^(2i/d) computed in log space for stability14 div_term = torch.exp(15 torch.arange(0, d_model, 2).float() * (-math.log(10000.0) / d_model)16 )1718 pe[:, 0::2] = torch.sin(position * div_term)19 pe[:, 1::2] = torch.cos(position * div_term)2021 # register_buffer: saved with the model, moved by .to(device), NOT a parameter22 self.register_buffer('pe', pe.unsqueeze(0)) # (1, max_len, d_model)2324 def forward(self, x): # (B, T, d_model)25 return self.dropout(x + self.pe[:, :x.size(1)])Two details that matter. register_buffer rather than nn.Parameter: the table is constant, so it should move with the model and appear in the state dict but never receive a gradient. And the log-space computation of div_term: it is the form the reference implementations use, and it gives the same numbers as computing 10000 ** (-2*i/d) directly (the smallest value is about 10−4, nowhere near float32's limits). What matters is the exponent: it must be 2i/d, and writing i/d is a common slip that squashes every frequency into a narrower band.
One more convention from the original paper: embeddings are multiplied by dmodel before the positional encoding is added. With dmodel=512 that is a factor of 22.6. The reason is scale matching — embedding weights initialised with variance 1/dmodel produce vectors far smaller than the positional signal, which ranges over [−1,1] in every coordinate. Skip the scaling and the position signal swamps the token identity early in training.
Choosing between them, and what came after
| Sinusoidal | Learned absolute | RoPE (rotary) | ALiBi | |
|---|---|---|---|---|
| Parameters | 0 | max_len×d | 0 | 0 |
| Where applied | Added to input embeddings | Added to input embeddings | Rotates Q and K inside every attention layer | Adds a distance penalty to attention scores |
| Encodes | Absolute, with relative structure available | Absolute only | Relative, exactly | Relative, as a linear recency bias |
| Beyond training length | Defined, but quality degrades | Impossible — no row exists | Degrades; extendable by interpolating frequencies | Extrapolates well by construction |
| Typical users | Original transformer | BERT, GPT-2, ViT | LLaMA, Mistral, most recent LLMs | BLOOM, some long-context models |
The practical guidance: if your sequences never exceed a known maximum and you have plenty of data, learned embeddings are simple and fine. If you need any length flexibility, use a relative scheme. RoPE has become the default for new language models because it gives exact relative positioning at zero parameter cost and interacts cleanly with the KV cache.
Part 2: making a deep stack trainable
The problem with depth
Stack 24 transformer layers naively and training fails. The gradient reaching layer 1 is the product of 24 Jacobians:
If each Jacobian scales the gradient by 0.9 on average, the factor reaching layer 1 is 0.924=0.080 — eight percent. At 0.8 it is 0.824=0.0047. At 48 layers with 0.9 it is 0.948=0.0064. The early layers barely move while the late layers train normally.
Residual connections
The fix is to change what a layer computes. Instead of
write
The layer now produces a correction to its input rather than a replacement. Differentiate:
The identity term is the whole point. Even if ∂F/∂x is tiny, the factor is close to I, not close to zero, so the product across 24 layers stays near 1 rather than decaying to 0.08. There is now a path from the loss to every layer that passes through no multiplications at all.
A second benefit, easy to overlook: a residual block can learn to do nothing. If F outputs zeros, the layer is the identity. That means adding layers can never make the model strictly less expressive, which is why very deep residual networks do not degrade the way very deep plain networks do.
Without residual connections a deep transformer does not train slowly — the early layers effectively do not train at all, and the model behaves like a shallow one with wasted parameters.
Layer normalisation
Residuals introduce their own problem: repeated addition makes activations grow. If each sub-layer's output has a scale comparable to its input, magnitudes compound layer over layer, and by layer 20 the activations are large enough to saturate nonlinearities and destabilise gradients.
Layer normalisation renormalises each token's vector independently:
where μ and σ2 are the mean and variance across the feature dimension of that one token, and γ,β∈Rd are learned scale and shift.
Work it through on x=[2,−1,4,3].
Mean: μ=(2−1+4+3)/4=8/4=2.
Deviations: [0,−3,2,1]. Squared: [0,9,4,1], sum 14. Variance σ2=14/4=3.5, so σ=1.8708.
Normalised (taking ϵ as negligible, γ=1, β=0):
Check the result: mean =(0−1.6036+1.0690+0.5345)/4≈0. Sum of squares =0+2.5715+1.1428+0.2857=4.0, so variance =4/4=1. Zero mean, unit variance, as promised.
The ϵ (typically 10−5) is not decoration. A token whose features are all equal has σ2=0 exactly, and without ϵ you divide by zero. This happens more often than you would expect — for instance on padding positions of a freshly initialised model.
Why not batch normalisation
| Batch norm | Layer norm | |
|---|---|---|
| Statistics computed over | The batch, per feature | The features, per token |
| Depends on other examples in the batch | Yes | No |
| Behaviour with variable-length padded sequences | Padding tokens pollute the batch statistics unless carefully excluded | Unaffected — each token is normalised alone |
| Batch size 1 at inference | Requires stored running statistics; train and eval behave differently | Identical in train and eval |
| Sequence length varies between batches | Statistics shift with the padding ratio | Irrelevant |
The decisive issue is padding. In a batch of sentences of lengths 8, 45 and 120, batch norm at position 100 would compute statistics over two padded positions and one real one — statistics that change with every batch composition. Layer norm never looks sideways, so none of this arises.
Add and Norm
Both mechanisms wrap every sub-layer. The original arrangement, post-norm:
Applied twice per encoder layer — once around attention, once around the feed-forward network.
Post-norm versus pre-norm
Look closely at post-norm and you will spot the flaw: the residual path is inside the normalisation. The gradient flowing backwards must pass through a LayerNorm at every layer, so the clean identity path is not clean after all. Empirically this makes deep post-norm transformers unstable without a learning-rate warmup — the loss diverges in the first few hundred steps.
Pre-norm moves the normalisation inside the residual branch:
Now the residual stream from input to output is a pure sum with no normalisation on it at all.
| Post-norm | Pre-norm | |
|---|---|---|
| Gradient path to layer 1 | Passes through L LayerNorms | Unobstructed sum |
| Warmup required | Yes — diverges without it | Optional; still helps |
| Stability at 24+ layers | Fragile | Robust |
| Final quality when both converge | Slightly better in several careful comparisons | Slightly worse |
| Extra requirement | None | A final LayerNorm after the last layer — the stream is otherwise never normalised on exit |
| Used by | Original transformer, BERT | GPT-2 onwards, LLaMA, most modern models |
The missing final LayerNorm in pre-norm is a real and easily-made bug. Without it the output magnitudes grow with depth and the logits come out badly scaled, producing a loss that starts far higher than ln(vocab size).
A complete encoder layer
1class EncoderLayer(nn.Module):2 """Supports both norm placements so you can compare them."""3 def __init__(self, d_model, num_heads, d_ff, dropout=0.1, pre_norm=True):4 super().__init__()5 self.pre_norm = pre_norm6 self.attn = MultiHeadAttention(d_model, num_heads, dropout)7 self.ff = FeedForward(d_model, d_ff, dropout)8 self.norm1 = nn.LayerNorm(d_model)9 self.norm2 = nn.LayerNorm(d_model)10 self.drop = nn.Dropout(dropout)1112 def forward(self, x, mask=None):13 if self.pre_norm:14 h = self.norm1(x)15 a, _ = self.attn(h, h, h, mask)16 x = x + self.drop(a)17 x = x + self.drop(self.ff(self.norm2(x)))18 else:19 a, _ = self.attn(x, x, x, mask)20 x = self.norm1(x + self.drop(a))21 x = self.norm2(x + self.drop(self.ff(x)))22 return x2324class Encoder(nn.Module):25 def __init__(self, num_layers, d_model, num_heads, d_ff,26 dropout=0.1, pre_norm=True):27 super().__init__()28 self.layers = nn.ModuleList(29 EncoderLayer(d_model, num_heads, d_ff, dropout, pre_norm)30 for _ in range(num_layers)31 )32 # required for pre-norm; harmless (an extra normalisation) for post-norm33 self.final_norm = nn.LayerNorm(d_model) if pre_norm else nn.Identity()3435 def forward(self, x, mask=None):36 for layer in self.layers:37 x = layer(x, mask)38 return self.final_norm(x)What this means when you build something
These three mechanisms fail in ways that look like different problems, so it is worth knowing their signatures.
| Symptom | Likely cause | Check |
|---|---|---|
| Model handles word identity fine but is insensitive to order — dog bites man and man bites dog score alike | Positional encoding not added, or added after the first layer | Feed a shuffled sequence and confirm the output is not a permutation of the original |
Loss explodes to nan in the first 200 steps of a deep model | Post-norm without warmup | Switch to pre-norm, or add a linear warmup over 4000 steps |
| Loss starts far above ln(V) and descends slowly | Pre-norm without the final LayerNorm, or missing dmodel embedding scale | A well-initialised classifier over V classes should start near ln(V): for V=32000, that is 10.4 |
| Early layers' weights barely change over training | Residual connections missing or applied only to some sub-layers | Log per-layer gradient norms; they should be within an order of magnitude of each other |
| Model fails on inputs longer than anything seen in training | Learned absolute position embeddings | Move to a relative scheme, or accept the hard length cap explicitly |
That third row is the cheapest sanity check in deep learning and almost nobody runs it. A classifier that starts at uniform probability over V classes has cross-entropy lnV: 6.9 for a 1000-class problem, 10.4 for a 32,000-token vocabulary. If your very first loss value is 25, something is wrong with initialisation or normalisation and no amount of training will be an efficient way to find out.
For new work the defaults are settled: pre-norm with a final LayerNorm, and a relative position scheme such as RoPE. Post-norm and sinusoidal encoding are not wrong — they are what the original models used and they work — but they carry requirements (warmup, a length ceiling) that you have to remember, and pre-norm plus RoPE do not.