Transformer Architecture Q&A

Course Content

Transformer Architecture Q&A

6 sections · 60 lessons

What is positional encoding, and why can a Transformer not function without position info?


What you need to know

Proof in five lines

Python
import numpy as nprng = np.random.default_rng(0)X = rng.normal(size=(3, 4))                    # "India beat Australia", no positionsW_q, W_k, W_v = (rng.normal(size=(4, 4)) for _ in range(3))def attend(X):    s = (X @ W_q) @ (X @ W_k).T / 2    w = np.exp(s - s.max(-1, keepdims=True))    w /= w.sum(-1, keepdims=True)    return w @ (X @ W_v)perm = [2, 1, 0]                               # "Australia beat India"print(np.allclose(attend(X)[perm], attend(X[perm])))   # True: same vectors, reordered

"India" gets the same output vector whether it comes first or last. The FFN works on each token alone, so it cannot recover order either. Without position information, the model reads a bag of words.

Sinusoidal encoding (2017)

The original paper added a fixed vector to each embedding:

Text
PE[pos, 2i]   = sin(pos / 10000^(2i / d))PE[pos, 2i+1] = cos(pos / 10000^(2i / d))

With d = 4, worked values (rounded):

Text
pos 0: [0.000,  1.000, 0.00, 1.00]pos 1: [0.841,  0.540, 0.01, 1.00]pos 2: [0.909, -0.416, 0.02, 1.00]pos 3: [0.141, -0.990, 0.03, 1.00]

Dimensions 0–1 change quickly (a "second hand"); dimensions 2–3 change slowly (an "hour hand"). Together they give each position a unique pattern. Moving by a fixed offset is a fixed rotation of each sin/cos pair, which was meant to help the model learn relative distances. No parameters are needed.

The main families

  • Absolute, added at the input — sinusoidal (original Transformer, Whisper's encoder) or learned vectors (GPT-2, BERT).
  • Relative, inside attention — RoPE rotates queries and keys (Llama, Qwen, Mistral, Gemma, DeepSeek); ALiBi adds a distance penalty to scores (BLOOM, MPT); T5 adds a learned bias per distance bucket.

A real-life example

An Indian-language news app translates cricket headlines into Hindi. "India beat Australia by 6 wickets" and "Australia beat India by 6 wickets" contain exactly the same tokens. Without positions, the encoder would produce the same set of vectors for both, and the decoder could not know who won — a very public mistake for a sports desk.

With positional information, "India" at position 1 and "India" at position 3 are different inputs, so attention can learn that the token before "beat" is the winner. The model then writes "भारत ने ऑस्ट्रेलिया को हराया" for one and the reverse for the other.

Follow-up questions to expect

  • "Does the causal mask give position information?" — Partly. Models trained with no positional encoding (NoPE) can infer some order from the mask, because each position sees a different number of tokens. It is weaker, so most models still add explicit positions.
  • "Why add positions instead of concatenating?" — Adding keeps d_model unchanged; in high dimensions the model can learn to keep content and position in nearly separate directions.
  • "Why sin and cos, not just the index?" — Raw indices grow without bound and would swamp the embedding; sinusoids stay between -1 and 1.