Course Content
Transformer Architecture Q&A
6 sections · 60 lessons
Explain RoPE and where in the attention pipeline it is applied?
What you need to know
The idea in two dimensions
Take a 2-D query and key. Rotate the query by m × θ (its position m) and the key by n × θ (its position n). A dot product between two rotated vectors depends only on the difference of the angles, so the score depends on m - n, not on m and n themselves.
1import numpy as np23def rotate(v, pos, theta=0.5):4 a = pos * theta # angle grows with position5 c, s = np.cos(a), np.sin(a)6 return np.array([c * v[0] - s * v[1], s * v[0] + c * v[1]])78q, k = np.array([1.0, 0.0]), np.array([0.6, 0.8])9for m, n in [(3, 1), (7, 5), (102, 100)]: # query pos m, key pos n10 print(m, n, round(rotate(q, m) @ rotate(k, n), 3))11# 3 1 0.99712# 7 5 0.99713# 102 100 0.997Three very different absolute positions, the same distance of 2, the same score. That is the whole trick: absolute rotations give a relative score.
In a real head
A head of size d has d / 2 pairs. Pair i rotates at its own frequency:
θ_i = base^(-2i / d) base = 10000 in the original; Llama 3 uses 500000angle for pair i at position m = m · θ_iq'_m = R(m θ) q_m, k'_n = R(n θ) k_nq'_m · k'_n depends only on (m - n)With d = 128 and base 10,000, pair 0 rotates 1 radian per token (it tracks nearby order), while the last pair rotates about 0.0001 radians per token (it changes slowly across thousands of tokens). A larger base slows all rotations, which helps with long contexts.
Where it sits
- Project —
q = x W_q,k = x W_k,v = x W_v. - Rotate — apply RoPE to
qandkusing each token's position; leavevalone. - Score —
q' k'ᵀ / sqrt(d_k), mask, softmax, multiply byv.
This happens in every layer, not once at the input. Keys are stored in the KV cache already rotated, so a cached key never needs recomputing.
Extending context
A model trained at 8k positions has never seen larger angles. Common fixes change the rotation so longer positions map into the familiar range: position interpolation (squeeze positions), NTK-aware scaling (raise the base), and YaRN (scale different frequencies differently, plus a temperature fix). They work best with a short fine-tune on long data. Some 2025 models, such as Llama 4, also mix in layers with no positional encoding to help long-range recall.
A real-life example
A code-completion assistant runs on a model trained with 8,192-token context. Users want completions that use code from other files in the repository, which needs about 32,000 tokens of context.
Simply feeding 32k tokens gives poor completions: positions beyond 8k produce rotation angles the model never trained on. The team applies YaRN-style scaling with a factor of 4 in the model config, then fine-tunes for a short run on long repository files. Completions that reference a helper function defined 20,000 tokens earlier now work. Because RoPE has no parameters, no weights had to be added or resized — only the rotation schedule changed.
Follow-up questions to expect
- "Why not rotate V as well?" — Position only needs to affect which tokens are matched (scores), not the content copied; rotating V would scramble the values.
- "Does RoPE work with the KV cache?" — Yes. Each key is rotated once with its own position and cached; new queries are rotated with theirs.
- "Is RoPE absolute or relative?" — It is applied using absolute positions, but the attention score depends only on relative distance.