Transformer Architecture Q&A

Course Content

Transformer Architecture Q&A

6 sections · 60 lessons

Explain RoPE and where in the attention pipeline it is applied?


Where RoPE sits inside every attention layerx from theresidual streamProjectto q, k, vRotate q and k byposition; v untouchedScore q'·k'— dependsonly on m − nMask, softmax,weights times vNo parameters; keys enter the KV cache already rotated.
Positions (3, 1), (7, 5) and (102, 100) all give the same score of 0.997 — absolute rotations produce relative attention.

What you need to know

The idea in two dimensions

Take a 2-D query and key. Rotate the query by m × θ (its position m) and the key by n × θ (its position n). A dot product between two rotated vectors depends only on the difference of the angles, so the score depends on m - n, not on m and n themselves.

Python
import numpy as npdef rotate(v, pos, theta=0.5):    a = pos * theta                            # angle grows with position    c, s = np.cos(a), np.sin(a)    return np.array([c * v[0] - s * v[1], s * v[0] + c * v[1]])q, k = np.array([1.0, 0.0]), np.array([0.6, 0.8])for m, n in [(3, 1), (7, 5), (102, 100)]:      # query pos m, key pos n    print(m, n, round(rotate(q, m) @ rotate(k, n), 3))# 3 1 0.997# 7 5 0.997# 102 100 0.997

Three very different absolute positions, the same distance of 2, the same score. That is the whole trick: absolute rotations give a relative score.

In a real head

A head of size d has d / 2 pairs. Pair i rotates at its own frequency:

Text
θ_i = base^(-2i / d)          base = 10000 in the original; Llama 3 uses 500000angle for pair i at position m = m · θ_iq'_m = R(m θ) q_m,   k'_n = R(n θ) k_nq'_m · k'_n  depends only on (m - n)

With d = 128 and base 10,000, pair 0 rotates 1 radian per token (it tracks nearby order), while the last pair rotates about 0.0001 radians per token (it changes slowly across thousands of tokens). A larger base slows all rotations, which helps with long contexts.

Where it sits

  1. Project — q = x W_q, k = x W_k, v = x W_v.
  2. Rotate — apply RoPE to q and k using each token's position; leave v alone.
  3. Score — q' k'ᵀ / sqrt(d_k), mask, softmax, multiply by v.

This happens in every layer, not once at the input. Keys are stored in the KV cache already rotated, so a cached key never needs recomputing.

Extending context

A model trained at 8k positions has never seen larger angles. Common fixes change the rotation so longer positions map into the familiar range: position interpolation (squeeze positions), NTK-aware scaling (raise the base), and YaRN (scale different frequencies differently, plus a temperature fix). They work best with a short fine-tune on long data. Some 2025 models, such as Llama 4, also mix in layers with no positional encoding to help long-range recall.

A real-life example

A code-completion assistant runs on a model trained with 8,192-token context. Users want completions that use code from other files in the repository, which needs about 32,000 tokens of context.

Simply feeding 32k tokens gives poor completions: positions beyond 8k produce rotation angles the model never trained on. The team applies YaRN-style scaling with a factor of 4 in the model config, then fine-tunes for a short run on long repository files. Completions that reference a helper function defined 20,000 tokens earlier now work. Because RoPE has no parameters, no weights had to be added or resized — only the rotation schedule changed.

Follow-up questions to expect

  • "Why not rotate V as well?" — Position only needs to affect which tokens are matched (scores), not the content copied; rotating V would scramble the values.
  • "Does RoPE work with the KV cache?" — Yes. Each key is rotated once with its own position and cached; new queries are rotated with theirs.
  • "Is RoPE absolute or relative?" — It is applied using absolute positions, but the attention score depends only on relative distance.