Transformer Architecture Q&A

Course Content

Transformer Architecture Q&A

6 sections · 60 lessons

Contrast ALiBi vs RoPE vs learned positional embeddings—what trade-offs do they represent?


ALiBi bias added to scores, query at position 40-0.5-1.0-1.5-2.00-0.004-0.008-0.012-0.016dist 0dist 1dist 2dist 3dist 4head 1, slope 1/2head 8, slope 1/256A fixed formula of distance, so it works at any length.
ALiBi's penalty only ever grows with distance, so it extrapolates cheaply — and always pushes the far token down, even when it holds the answer.

What you need to know

Learned absoluteRoPEALiBi
Where appliedAdded to embeddings at inputRotates Q and K in every layerAdds a bias to the score matrix
Position typeAbsoluteRelative (via absolute rotation)Relative (raw distance)
Parametersmax_len × d_model, trainedNoneNone (fixed slope per head)
Beyond trained lengthImpossible: no vector existsDegrades; extendable with scaling plus fine-tuneExtrapolates gracefully
Used byGPT-2, BERT, Whisper decoderLlama, Qwen, Mistral, Gemma, DeepSeekBLOOM, MPT

Learned absolute

A table of shape (max_len, d_model), just like the token embedding table. GPT-2 has 1,024 rows, BERT 512. Position 1,025 simply has no vector. Rare late positions also get less training, so they are learned less well.

ALiBi, worked by hand

ALiBi adds -m × distance to each score, where m is a slope fixed per head. With 8 heads the slopes are 1/2, 1/4, 1/8, ... 1/256. For a query at position 4 looking back:

Text
distance to keys 4, 3, 2, 1, 0 :   0     1     2     3     4bias, head 1 (m = 1/2)         :   0   -0.5  -1.0  -1.5  -2.0bias, head 8 (m = 1/256)       :   0   -0.004 -0.008 -0.012 -0.016

Head 1 strongly prefers nearby tokens; head 8 barely cares about distance. Because the bias is just a formula of distance, it works at any length. The cost: the penalty only grows, so a key 10,000 tokens away is always pushed down, even when it holds the answer.

RoPE

No parameters and relative by design. Its long-context tools (position interpolation, NTK scaling, YaRN) are well tested and supported in common inference servers. The trade-off is that extension usually needs a short fine-tune to keep quality.

A real-life example

A speech-to-text team uses Whisper. Its encoder uses fixed sinusoidal positions over 1,500 audio frames (30 seconds), and its decoder uses learned positions with a maximum of 448 text tokens. An engineer tries to transcribe a two-minute audio clip in one pass and finds it cannot work: there are no encoder positions past 30 seconds and no decoder positions past 448 tokens. The standard pipeline instead cuts audio into 30-second chunks.

A few months later the same team designs its own long-form transcription model. They pick RoPE for the decoder: no position table to cap length, cache-friendly, and extendable later with scaling. They rule out ALiBi because they need the model to recall a customer's name from early in a long call, and a distance penalty works against that. The choice follows directly from the trade-offs in the table.

Follow-up questions to expect

  • "Which extrapolates best without fine-tuning?" — ALiBi, by design. RoPE needs scaling; learned embeddings cannot extrapolate at all.
  • "What does T5 use?" — A learned bias per bucket of relative distance, added to the scores — a close cousin of ALiBi, but trained.
  • "Can you mix schemes?" — Yes. Some recent models interleave RoPE layers with no-position (NoPE) layers to improve long-range recall.