Course Content
LLMs Deep Dive
10 sections · 40 lessons
What are positional encodings in LLMs?
What you need to know
Why attention needs help
Attention scores come from dot products between token vectors. If you shuffle the tokens, every token still has the same vector and gets the same scores — only the order of rows changes. So "merchant paid customer" and "customer paid merchant" produce the same set of outputs, even though one is a refund and the other a sale.
The main approaches
| Method | How it works | Used in |
|---|---|---|
| Sinusoidal | Add fixed sine/cosine waves of different frequencies to each embedding | Original 2017 transformer |
| Learned absolute | Learn one vector per position, 0 to max length | BERT, GPT-2 |
| RoPE (rotary) | Rotate query and key vectors by position-dependent angles | Llama, Qwen, Mistral, DeepSeek, most modern LLMs |
| ALiBi | Subtract a penalty proportional to distance from each attention score | BLOOM, MPT |
Learned absolute positions have a hard limit: a model trained with 1,024 positions has no vector for position 1,025.
Why RoPE gives relative position
Take one pair of dimensions and a frequency of 10 degrees per position. A query at position 7 is rotated by 70°. A key at position 5 is rotated by 50°. The dot product between two rotated vectors depends only on the difference in angles, 20°, which corresponds to a distance of 2. Move both tokens 10 positions later, to positions 17 and 15 (170° and 150°), and the difference is still 20°. So attention sees "2 tokens apart", not "position 7 and position 5". Different pairs use different frequencies, so some pairs track short distances and others long ones.
Extending context
A model trained on 8,000-token sequences has only seen angles up to a certain size. Show it position 100,000 and the angles are new, so quality collapses. Methods such as position interpolation, NTK-aware scaling and YaRN rescale the RoPE frequencies so long positions map into the familiar range, followed by a short fine-tune on long documents. This is how many open models went from 8K to 128K tokens without training from scratch.
A real-life example
A bank's chatbot handles refund questions in Hinglish: "Merchant ne mujhe 500 bheje, ya maine merchant ko?" ("Did the merchant send me 500, or did I send it to the merchant?"). Getting direction right decides whether the bot says "you will receive a refund" or "your payment was successful".
Positional information is what lets attention link "bheje" (sent) with the correct sender and receiver based on order. Separately, the same bank tried an open model on 60-page loan agreements longer than its trained context. Answers about the last pages were garbled. Switching to a version whose RoPE had been scaled and fine-tuned for 128K tokens fixed it — the fix was in position handling, not in the prompt.
Follow-up questions to expect
- "Why not just use learned positions?" — They cannot go past the maximum length seen in training, and they encode absolute rather than relative position.
- "Does RoPE change the embeddings?" — No; it rotates queries and keys inside every attention layer, so values and the residual stream carry no position vector.
- "What is the 'lost in the middle' effect, and is it about positions?" — Models recall information at the start and end of long contexts better than in the middle. Position handling and training data both contribute to it.