Course Content
Transformer Architecture Q&A
6 sections · 60 lessons
Contrast ALiBi vs RoPE vs learned positional embeddings—what trade-offs do they represent?
What you need to know
| Learned absolute | RoPE | ALiBi | |
|---|---|---|---|
| Where applied | Added to embeddings at input | Rotates Q and K in every layer | Adds a bias to the score matrix |
| Position type | Absolute | Relative (via absolute rotation) | Relative (raw distance) |
| Parameters | max_len × d_model, trained | None | None (fixed slope per head) |
| Beyond trained length | Impossible: no vector exists | Degrades; extendable with scaling plus fine-tune | Extrapolates gracefully |
| Used by | GPT-2, BERT, Whisper decoder | Llama, Qwen, Mistral, Gemma, DeepSeek | BLOOM, MPT |
Learned absolute
A table of shape (max_len, d_model), just like the token embedding table. GPT-2 has 1,024 rows, BERT 512. Position 1,025 simply has no vector. Rare late positions also get less training, so they are learned less well.
ALiBi, worked by hand
ALiBi adds -m × distance to each score, where m is a slope fixed per head. With 8 heads the slopes are 1/2, 1/4, 1/8, ... 1/256. For a query at position 4 looking back:
distance to keys 4, 3, 2, 1, 0 : 0 1 2 3 4bias, head 1 (m = 1/2) : 0 -0.5 -1.0 -1.5 -2.0bias, head 8 (m = 1/256) : 0 -0.004 -0.008 -0.012 -0.016Head 1 strongly prefers nearby tokens; head 8 barely cares about distance. Because the bias is just a formula of distance, it works at any length. The cost: the penalty only grows, so a key 10,000 tokens away is always pushed down, even when it holds the answer.
RoPE
No parameters and relative by design. Its long-context tools (position interpolation, NTK scaling, YaRN) are well tested and supported in common inference servers. The trade-off is that extension usually needs a short fine-tune to keep quality.
A real-life example
A speech-to-text team uses Whisper. Its encoder uses fixed sinusoidal positions over 1,500 audio frames (30 seconds), and its decoder uses learned positions with a maximum of 448 text tokens. An engineer tries to transcribe a two-minute audio clip in one pass and finds it cannot work: there are no encoder positions past 30 seconds and no decoder positions past 448 tokens. The standard pipeline instead cuts audio into 30-second chunks.
A few months later the same team designs its own long-form transcription model. They pick RoPE for the decoder: no position table to cap length, cache-friendly, and extendable later with scaling. They rule out ALiBi because they need the model to recall a customer's name from early in a long call, and a distance penalty works against that. The choice follows directly from the trade-offs in the table.
Follow-up questions to expect
- "Which extrapolates best without fine-tuning?" — ALiBi, by design. RoPE needs scaling; learned embeddings cannot extrapolate at all.
- "What does T5 use?" — A learned bias per bucket of relative distance, added to the scores — a close cousin of ALiBi, but trained.
- "Can you mix schemes?" — Yes. Some recent models interleave RoPE layers with no-position (NoPE) layers to improve long-range recall.