Attention Mechanisms and Transformers

Positional Encoding & Residual Connections


Here is an experiment you can run in five minutes. Build a self-attention layer, feed it the token embeddings for dog bites man, and record the output. Now shuffle the input rows to man bites dog and run it again. Compare.

The outputs are identical — just permuted the same way the input was. Not similar: bit-for-bit the same vectors, in a different order. The layer has no idea the order changed, because nothing in softmax(QK⊤/dk)V\mathrm{softmax}(QK^{\top}/\sqrt{d_k})V references a position index. Scores depend only on which vectors are present, so a permutation of the input produces exactly the corresponding permutation of the output.

That property has a name — permutation equivariance — and it makes raw self-attention a very expensive bag of words. A model built this way will get word identity right and word order consistently wrong, and no amount of extra layers will fix it, because stacking permutation-equivariant layers gives you another permutation-equivariant function.

There are two separate structural problems to solve before a stack of attention layers becomes a working transformer. The first is telling the model where each token is. The second is making a deep stack of these layers trainable at all. This lesson does both, with the arithmetic.

Attention sees a set, so position has to be added indog+ p0bites+ p1man+ p2012subjectobjectShuffle the rows without the p terms and every output row is unchanged, only reordered.
Position is added, not concatenated, so every dimension carries both meaning and place at no extra width.

Part 1: injecting position

Why it must be added to the input

Since attention cannot see position, position must be baked into the vectors it does see. The standard approach adds a position-dependent vector to each token embedding before the first layer:

xi=Embed(tokeni)+PE(i)x_i = \mathrm{Embed}(\text{token}_i) + \mathrm{PE}(i)

Adding rather than concatenating is a deliberate choice. Concatenation would preserve the two signals in separate coordinates but widen every matrix in the model. Addition keeps dmodeld_{model} fixed and relies on the model learning to separate the two contributions — which it can, because in high dimensions two random subspaces are nearly orthogonal, and because the projections WQW^Q and WKW^K are free to attend to whichever coordinates carry which signal.

Option A: learned absolute embeddings

The simplest thing that works: a lookup table with one trainable vector per position.

Python
class LearnedPositionalEmbedding(nn.Module):    def __init__(self, max_len, d_model):        super().__init__()        self.pos_emb = nn.Embedding(max_len, d_model)    def forward(self, x):                      # x: (B, T, d_model)        T = x.size(1)        positions = torch.arange(T, device=x.device)        return x + self.pos_emb(positions)     # broadcasts over the batch

Parameter cost: max_len × d_model. For BERT-base that is 512×768=393,216512 \times 768 = 393{,}216 parameters — about 0.36% of its 110 M total, which is negligible.

The hard limit is in the shape of that table. Position 512 has a row; position 513 does not exist. Feed a longer sequence and you get an index error, or worse, a silent wrap-around if someone wrote a modulo. There is no principled way to extrapolate, because position 513's embedding was never trained and there is no formula connecting it to position 512's.

Option B: sinusoidal encoding

The original transformer used a fixed formula instead. For position pospos and dimension index ii:

PE(pos,2i)=sin⁡ ⁣(pos100002i/dmodel),PE(pos,2i+1)=cos⁡ ⁣(pos100002i/dmodel)\mathrm{PE}(pos, 2i) = \sin\!\left(\frac{pos}{10000^{2i/d_{model}}}\right), \qquad \mathrm{PE}(pos, 2i+1) = \cos\!\left(\frac{pos}{10000^{2i/d_{model}}}\right)

Dimensions are paired: each pair (2i,2i+1)(2i, 2i+1) is a sine and cosine at one frequency, and the frequency decreases geometrically as ii increases. Pair 0 oscillates with wavelength 2π≈6.32\pi \approx 6.3 positions; the last pair has wavelength 2π⋅10000≈62,8322\pi \cdot 10000 \approx 62{,}832 positions. Together they act like a clock face with hands of many different speeds — the fast hands distinguish neighbours, the slow hands distinguish regions of the sequence.

Compute it by hand for dmodel=4d_{model} = 4. There are two pairs. For pair i=0i=0 the divisor is 100000=110000^{0} = 1; for pair i=1i=1 it is 100002/4=100000.5=10010000^{2/4} = 10000^{0.5} = 100.

posdim 0: sin⁡(pos)\sin(pos)dim 1: cos⁡(pos)\cos(pos)dim 2: sin⁡(pos/100)\sin(pos/100)dim 3: cos⁡(pos/100)\cos(pos/100)
00.00001.00000.00001.00000
10.84150.54030.01000.99995
20.9093−0.41610.02000.99980
30.1411−0.99000.03000.99955

Every row is different, so positions are distinguishable. But the interesting property is what happens when you take dot products between rows.

PE(1)⋅PE(2)=(0.8415)(0.9093)+(0.5403)(−0.4161)+(0.0100)(0.0200)+(0.99995)(0.99980)\mathrm{PE}(1)\cdot\mathrm{PE}(2) = (0.8415)(0.9093) + (0.5403)(-0.4161) + (0.0100)(0.0200) + (0.99995)(0.99980)

=0.76518−0.22482+0.00020+0.99975=1.5403= 0.76518 - 0.22482 + 0.00020 + 0.99975 = 1.5403
PE(2)⋅PE(3)=(0.9093)(0.1411)+(−0.4161)(−0.9900)+(0.0200)(0.0300)+(0.99980)(0.99955)\mathrm{PE}(2)\cdot\mathrm{PE}(3) = (0.9093)(0.1411) + (-0.4161)(-0.9900) + (0.0200)(0.0300) + (0.99980)(0.99955)
=0.12830+0.41194+0.00060+0.99935=1.5402= 0.12830 + 0.41194 + 0.00060 + 0.99935 = 1.5402

The same value. Positions 1-and-2 and positions 2-and-3 have identical similarity, because both pairs are one step apart. This is not a coincidence — it falls out of the angle-difference identity:

sin⁡asin⁡b+cos⁡acos⁡b=cos⁡(a−b)\sin a \sin b + \cos a \cos b = \cos(a - b)

so the dot product over all pairs is ∑icos⁡ ⁣(Δ100002i/d)\sum_i \cos\!\left(\frac{\Delta}{10000^{2i/d}}\right), a function of the offset Δ\Delta alone. For Δ=1\Delta = 1 here that is cos⁡(1)+cos⁡(0.01)=0.5403+0.99995=1.54025\cos(1) + \cos(0.01) = 0.5403 + 0.99995 = 1.54025, matching both computations above.

There is a stronger property. PE(pos+k)\mathrm{PE}(pos + k) is a fixed linear function of PE(pos)\mathrm{PE}(pos) — the same function for every pospos. Within one frequency pair, write the rotation matrix

Rk=[cos⁡ksin⁡k−sin⁡kcos⁡k]R_k = \begin{bmatrix} \cos k & \sin k \\ -\sin k & \cos k \end{bmatrix}

Then

Rk[sin⁡poscos⁡pos]=[cos⁡ksin⁡pos+sin⁡kcos⁡pos−sin⁡ksin⁡pos+cos⁡kcos⁡pos]=[sin⁡(pos+k)cos⁡(pos+k)]R_k \begin{bmatrix}\sin pos \\ \cos pos\end{bmatrix} = \begin{bmatrix} \cos k \sin pos + \sin k \cos pos \\ -\sin k \sin pos + \cos k \cos pos \end{bmatrix} = \begin{bmatrix}\sin(pos+k) \\ \cos(pos+k)\end{bmatrix}

Because WQW^Q and WKW^K are linear maps, the model can in principle learn to implement "attend to whatever is 3 positions back" as a single linear operation that works at every position. That is what makes sinusoidal encoding more than an arbitrary set of distinct labels.

Sinusoidal encoding gives every pair of positions a similarity that depends only on their distance, and makes relative offsets expressible as fixed linear maps — which is why a formula beats a lookup table despite having zero parameters.

Python
import mathimport torchimport torch.nn as nnclass SinusoidalPositionalEncoding(nn.Module):    def __init__(self, d_model, max_len=5000, dropout=0.1):        super().__init__()        self.dropout = nn.Dropout(dropout)        pe = torch.zeros(max_len, d_model)        position = torch.arange(max_len).unsqueeze(1).float()      # (max_len, 1)        # 1 / 10000^(2i/d) computed in log space for stability        div_term = torch.exp(            torch.arange(0, d_model, 2).float() * (-math.log(10000.0) / d_model)        )        pe[:, 0::2] = torch.sin(position * div_term)        pe[:, 1::2] = torch.cos(position * div_term)        # register_buffer: saved with the model, moved by .to(device), NOT a parameter        self.register_buffer('pe', pe.unsqueeze(0))                # (1, max_len, d_model)    def forward(self, x):                                          # (B, T, d_model)        return self.dropout(x + self.pe[:, :x.size(1)])

Two details that matter. register_buffer rather than nn.Parameter: the table is constant, so it should move with the model and appear in the state dict but never receive a gradient. And the log-space computation of div_term: it is the form the reference implementations use, and it gives the same numbers as computing 10000 ** (-2*i/d) directly (the smallest value is about 10−410^{-4}, nowhere near float32's limits). What matters is the exponent: it must be 2i/d2i/d, and writing i/di/d is a common slip that squashes every frequency into a narrower band.

One more convention from the original paper: embeddings are multiplied by dmodel\sqrt{d_{model}} before the positional encoding is added. With dmodel=512d_{model}=512 that is a factor of 22.6. The reason is scale matching — embedding weights initialised with variance 1/dmodel1/d_{model} produce vectors far smaller than the positional signal, which ranges over [−1,1][-1, 1] in every coordinate. Skip the scaling and the position signal swamps the token identity early in training.

Choosing between them, and what came after

SinusoidalLearned absoluteRoPE (rotary)ALiBi
Parameters0max_len×d\text{max\_len} \times d00
Where appliedAdded to input embeddingsAdded to input embeddingsRotates Q and K inside every attention layerAdds a distance penalty to attention scores
EncodesAbsolute, with relative structure availableAbsolute onlyRelative, exactlyRelative, as a linear recency bias
Beyond training lengthDefined, but quality degradesImpossible — no row existsDegrades; extendable by interpolating frequenciesExtrapolates well by construction
Typical usersOriginal transformerBERT, GPT-2, ViTLLaMA, Mistral, most recent LLMsBLOOM, some long-context models

The practical guidance: if your sequences never exceed a known maximum and you have plenty of data, learned embeddings are simple and fine. If you need any length flexibility, use a relative scheme. RoPE has become the default for new language models because it gives exact relative positioning at zero parameter cost and interacts cleanly with the KV cache.

Part 2: making a deep stack trainable

The problem with depth

Stack 24 transformer layers naively and training fails. The gradient reaching layer 1 is the product of 24 Jacobians:

∂L∂x1=∂L∂x24∏ℓ=224∂xℓ∂xℓ−1\frac{\partial \mathcal{L}}{\partial x_1} = \frac{\partial \mathcal{L}}{\partial x_{24}} \prod_{\ell=2}^{24}\frac{\partial x_\ell}{\partial x_{\ell-1}}

If each Jacobian scales the gradient by 0.9 on average, the factor reaching layer 1 is 0.924=0.0800.9^{24} = 0.080 — eight percent. At 0.8 it is 0.824=0.00470.8^{24} = 0.0047. At 48 layers with 0.9 it is 0.948=0.00640.9^{48} = 0.0064. The early layers barely move while the late layers train normally.

Residual connections

The fix is to change what a layer computes. Instead of

xℓ=F(xℓ−1)x_{\ell} = F(x_{\ell-1})

write

xℓ=xℓ−1+F(xℓ−1)x_{\ell} = x_{\ell-1} + F(x_{\ell-1})

The layer now produces a correction to its input rather than a replacement. Differentiate:

∂xℓ∂xℓ−1=I+∂F∂xℓ−1\frac{\partial x_\ell}{\partial x_{\ell-1}} = I + \frac{\partial F}{\partial x_{\ell-1}}

The identity term is the whole point. Even if ∂F/∂x\partial F/\partial x is tiny, the factor is close to II, not close to zero, so the product across 24 layers stays near 1 rather than decaying to 0.08. There is now a path from the loss to every layer that passes through no multiplications at all.

A second benefit, easy to overlook: a residual block can learn to do nothing. If FF outputs zeros, the layer is the identity. That means adding layers can never make the model strictly less expressive, which is why very deep residual networks do not degrade the way very deep plain networks do.

Without residual connections a deep transformer does not train slowly — the early layers effectively do not train at all, and the model behaves like a shallow one with wasted parameters.

Layer normalisation

Residuals introduce their own problem: repeated addition makes activations grow. If each sub-layer's output has a scale comparable to its input, magnitudes compound layer over layer, and by layer 20 the activations are large enough to saturate nonlinearities and destabilise gradients.

Layer normalisation renormalises each token's vector independently:

LN(x)=γ⊙x−μσ2+ϵ+β\mathrm{LN}(x) = \gamma \odot \frac{x - \mu}{\sqrt{\sigma^2 + \epsilon}} + \beta

where μ\mu and σ2\sigma^2 are the mean and variance across the feature dimension of that one token, and γ,β∈Rd\gamma, \beta \in \mathbb{R}^{d} are learned scale and shift.

Work it through on x=[2,−1,4,3]x = [2, -1, 4, 3].

Mean: μ=(2−1+4+3)/4=8/4=2\mu = (2 - 1 + 4 + 3)/4 = 8/4 = 2.

Deviations: [0,−3,2,1][0, -3, 2, 1]. Squared: [0,9,4,1][0, 9, 4, 1], sum 14. Variance σ2=14/4=3.5\sigma^2 = 14/4 = 3.5, so σ=1.8708\sigma = 1.8708.

Normalised (taking ϵ\epsilon as negligible, γ=1\gamma = 1, β=0\beta = 0):

[01.8708, −31.8708, 21.8708, 11.8708]=[0, −1.6036, 1.0690, 0.5345]\left[\frac{0}{1.8708},\ \frac{-3}{1.8708},\ \frac{2}{1.8708},\ \frac{1}{1.8708}\right] = [0,\ -1.6036,\ 1.0690,\ 0.5345]

Check the result: mean =(0−1.6036+1.0690+0.5345)/4≈0= (0 - 1.6036 + 1.0690 + 0.5345)/4 \approx 0. Sum of squares =0+2.5715+1.1428+0.2857=4.0= 0 + 2.5715 + 1.1428 + 0.2857 = 4.0, so variance =4/4=1= 4/4 = 1. Zero mean, unit variance, as promised.

The ϵ\epsilon (typically 10−510^{-5}) is not decoration. A token whose features are all equal has σ2=0\sigma^2 = 0 exactly, and without ϵ\epsilon you divide by zero. This happens more often than you would expect — for instance on padding positions of a freshly initialised model.

Why not batch normalisation

Batch normLayer norm
Statistics computed overThe batch, per featureThe features, per token
Depends on other examples in the batchYesNo
Behaviour with variable-length padded sequencesPadding tokens pollute the batch statistics unless carefully excludedUnaffected — each token is normalised alone
Batch size 1 at inferenceRequires stored running statistics; train and eval behave differentlyIdentical in train and eval
Sequence length varies between batchesStatistics shift with the padding ratioIrrelevant

The decisive issue is padding. In a batch of sentences of lengths 8, 45 and 120, batch norm at position 100 would compute statistics over two padded positions and one real one — statistics that change with every batch composition. Layer norm never looks sideways, so none of this arises.

Add and Norm

Both mechanisms wrap every sub-layer. The original arrangement, post-norm:

xℓ=LayerNorm(xℓ−1+Dropout(SubLayer(xℓ−1)))x_{\ell} = \mathrm{LayerNorm}\big(x_{\ell-1} + \mathrm{Dropout}(\mathrm{SubLayer}(x_{\ell-1}))\big)

Applied twice per encoder layer — once around attention, once around the feed-forward network.

Post-norm versus pre-norm

Look closely at post-norm and you will spot the flaw: the residual path is inside the normalisation. The gradient flowing backwards must pass through a LayerNorm at every layer, so the clean identity path is not clean after all. Empirically this makes deep post-norm transformers unstable without a learning-rate warmup — the loss diverges in the first few hundred steps.

Pre-norm moves the normalisation inside the residual branch:

xℓ=xℓ−1+Dropout(SubLayer(LayerNorm(xℓ−1)))x_{\ell} = x_{\ell-1} + \mathrm{Dropout}\big(\mathrm{SubLayer}(\mathrm{LayerNorm}(x_{\ell-1}))\big)

Now the residual stream from input to output is a pure sum with no normalisation on it at all.

Post-normPre-norm
Gradient path to layer 1Passes through LL LayerNormsUnobstructed sum
Warmup requiredYes — diverges without itOptional; still helps
Stability at 24+ layersFragileRobust
Final quality when both convergeSlightly better in several careful comparisonsSlightly worse
Extra requirementNoneA final LayerNorm after the last layer — the stream is otherwise never normalised on exit
Used byOriginal transformer, BERTGPT-2 onwards, LLaMA, most modern models

The missing final LayerNorm in pre-norm is a real and easily-made bug. Without it the output magnitudes grow with depth and the logits come out badly scaled, producing a loss that starts far higher than ln⁡(vocab size)\ln(\text{vocab size}).

A complete encoder layer

Python
class EncoderLayer(nn.Module):    """Supports both norm placements so you can compare them."""    def __init__(self, d_model, num_heads, d_ff, dropout=0.1, pre_norm=True):        super().__init__()        self.pre_norm = pre_norm        self.attn  = MultiHeadAttention(d_model, num_heads, dropout)        self.ff    = FeedForward(d_model, d_ff, dropout)        self.norm1 = nn.LayerNorm(d_model)        self.norm2 = nn.LayerNorm(d_model)        self.drop  = nn.Dropout(dropout)    def forward(self, x, mask=None):        if self.pre_norm:            h = self.norm1(x)            a, _ = self.attn(h, h, h, mask)            x = x + self.drop(a)            x = x + self.drop(self.ff(self.norm2(x)))        else:            a, _ = self.attn(x, x, x, mask)            x = self.norm1(x + self.drop(a))            x = self.norm2(x + self.drop(self.ff(x)))        return xclass Encoder(nn.Module):    def __init__(self, num_layers, d_model, num_heads, d_ff,                 dropout=0.1, pre_norm=True):        super().__init__()        self.layers = nn.ModuleList(            EncoderLayer(d_model, num_heads, d_ff, dropout, pre_norm)            for _ in range(num_layers)        )        # required for pre-norm; harmless (an extra normalisation) for post-norm        self.final_norm = nn.LayerNorm(d_model) if pre_norm else nn.Identity()    def forward(self, x, mask=None):        for layer in self.layers:            x = layer(x, mask)        return self.final_norm(x)

What this means when you build something

These three mechanisms fail in ways that look like different problems, so it is worth knowing their signatures.

SymptomLikely causeCheck
Model handles word identity fine but is insensitive to order — dog bites man and man bites dog score alikePositional encoding not added, or added after the first layerFeed a shuffled sequence and confirm the output is not a permutation of the original
Loss explodes to nan in the first 200 steps of a deep modelPost-norm without warmupSwitch to pre-norm, or add a linear warmup over 4000 steps
Loss starts far above ln⁡(V)\ln(V) and descends slowlyPre-norm without the final LayerNorm, or missing dmodel\sqrt{d_{model}} embedding scaleA well-initialised classifier over VV classes should start near ln⁡(V)\ln(V): for V=32000V = 32000, that is 10.4
Early layers' weights barely change over trainingResidual connections missing or applied only to some sub-layersLog per-layer gradient norms; they should be within an order of magnitude of each other
Model fails on inputs longer than anything seen in trainingLearned absolute position embeddingsMove to a relative scheme, or accept the hard length cap explicitly

That third row is the cheapest sanity check in deep learning and almost nobody runs it. A classifier that starts at uniform probability over VV classes has cross-entropy ln⁡V\ln V: 6.9 for a 1000-class problem, 10.4 for a 32,000-token vocabulary. If your very first loss value is 25, something is wrong with initialisation or normalisation and no amount of training will be an efficient way to find out.

For new work the defaults are settled: pre-norm with a final LayerNorm, and a relative position scheme such as RoPE. Post-norm and sinusoidal encoding are not wrong — they are what the original models used and they work — but they carry requirements (warmup, a length ceiling) that you have to remember, and pre-norm plus RoPE do not.