Course Content
How Large Language Models Work
3 sections · 9 lessons
Embeddings & Positional Encoding
A tokeniser has just handed the model a list of integers: [464, 3797, 3332]. Those numbers stand for The, cat, sat. They are IDs, nothing more — assigned in the order the tokeniser happened to learn them.
Feed those integers straight into a neural network and you have told it something false. Arithmetic on IDs implies relationships that do not exist. Token 3797 is not "larger" than token 464. The average of 464 and 3332 is 1898, which is some unrelated token. Whatever the network computes from these numbers, it will be computing over an accidental ordering.
So the very first learned layer of every language model exists to fix this: it replaces each arbitrary integer with a vector — a list of real numbers in which distance actually means something. And once that is done, a second problem appears immediately, one that is easy to miss and impossible to ignore: the core of a transformer has no idea what order its inputs arrived in. Both problems, and their solutions, are the subject here.
From integer to vector
The obvious first fix is one-hot encoding. Represent token 3797 as a vector of length 50,257 that is 0 everywhere except a single 1 at index 3797.
This removes the false ordering — every token is now exactly the same distance from every other. But it replaces one problem with another. Every pair of one-hot vectors has a dot product of exactly 0. cat and dog are as unrelated as cat and photosynthesis. The representation is honest about knowing nothing, and it is also 50,257 numbers wide, of which 50,256 are zero.
The fix is an embedding matrix: a learned table E with one row per vocabulary entry and dmodel columns. Looking up token t means taking row t.
Here is a detail worth internalising, because it explains why this is a normal, trainable layer rather than a special case: multiplying a one-hot vector by a matrix is a row lookup. With a 4-token vocabulary and dmodel=3:
one-hot(2) E (4 x 3) result[0 0 1 0] x [ 0.1 -0.3 0.8 ] = [ 0.5 0.9 -0.2 ] [ 0.4 0.2 -0.6 ] (row 2, exactly) [ 0.5 0.9 -0.2 ] [-0.7 0.1 0.3 ]So the lookup table is a linear layer, gradients flow through it, and the vectors are learned like any other weights. In code it is implemented as an indexing operation purely because indexing is faster than multiplying by a mostly-zero matrix.
What the learned vectors end up meaning
Nobody assigns meaning to the dimensions. Meaning emerges because the training objective punishes bad representations. If cat and dog get very different vectors, the model does badly at predicting the words that follow both of them, so gradient descent nudges them together. Words that appear in similar contexts drift towards similar vectors — not because anyone taught the model about animals, but because the loss goes down when it happens.
You can measure the result with cosine similarity, which asks how aligned two vectors are regardless of their length:
Take three toy 4-dimensional embeddings and do the arithmetic:
cat = [ 0.8, 0.6, 0.1, -0.2]dog = [ 0.7, 0.7, 0.2, -0.1]car = [-0.5, 0.3, 0.9, 0.4]cat . dog = (0.8)(0.7) + (0.6)(0.7) + (0.1)(0.2) + (-0.2)(-0.1) = 0.56 + 0.42 + 0.02 + 0.02 = 1.02|cat| = sqrt(0.64+0.36+0.01+0.04) = sqrt(1.05) = 1.0247|dog| = sqrt(0.49+0.49+0.04+0.01) = sqrt(1.03) = 1.0149cos(cat, dog) = 1.02 / (1.0247 x 1.0149) = 0.981cat . car = (0.8)(-0.5) + (0.6)(0.3) + (0.1)(0.9) + (-0.2)(0.4) = -0.40 + 0.18 + 0.09 - 0.08 = -0.21|car| = sqrt(0.25+0.09+0.81+0.16) = sqrt(1.31) = 1.1446cos(cat, car) = -0.21 / (1.0247 x 1.1446) = -0.1790.981 against −0.179. Under one-hot encoding both numbers would have been exactly 0. That difference — a geometry in which similarity is measurable — is what the embedding layer buys.
The embedding matrix is not a dictionary someone wrote. It is a map that the training process draws, where the only rule is: things that behave alike in context should sit near each other.
Why embeddings are usually scaled by dmodel
Many implementations multiply the looked-up embedding by dmodel before anything else touches it. This is not decoration. Embedding weights are typically initialised from a distribution with a small variance (often around 1/dmodel), which makes the raw vectors small. Positional information is about to be added on top with a magnitude of roughly 1 per component. Without the rescale, position would drown out identity at the very first layer. With dmodel=512, the factor is 512≈22.6 — enough to put the two signals on comparable footing.
The order problem
Now the second problem, which is easier to state than to believe until you check it.
The engine of a transformer is self-attention. Every position produces a query, a key and a value; every position's output is a weighted average of all the values, weighted by how well its query matches each key. Written out for position i:
Look hard at that sum. It runs over j. There is no j anywhere else in the expression — no term that depends on where token j sits. It is a sum over an unordered collection.
The consequence: shuffle the input tokens and the output vectors are shuffled identically, but otherwise unchanged. Consider two sentences built from the same three tokens:
TextSentence A: dog bites manSentence B: man bites dogThe multiset of (query, key, value) triples is identical in both. When the model computes the output at the position holding bites, it forms exactly the same weighted average in both cases, because it is averaging over the same set. To a position-free transformer, "dog bites man" and "man bites dog" are the same input. So are "not good" and "good not", and "3 minus 5" and "5 minus 3".
The obvious fix, and why it fails
Just append the position number to each vector. Token at position 0 gets an extra component 0, position 1 gets 1, and so on.
This breaks in three separate ways, and each is instructive:
- Unbounded magnitude. At position 4,000 that component is 4,000, while every other component sits near ±1. It dominates every dot product. The model attends on position and nothing else.
- No generalisation past training length. If training saw positions 0–1023, the value 4,000 is outside anything the weights ever adapted to. Behaviour there is undefined.
- Wrong quantity encoded. Language cares far more about relative distance than absolute index. "The adjective two tokens back" is a useful pattern; "the token at absolute index 743" almost never is. A raw index makes relative distance something the model must reconstruct by subtraction, at every layer, on its own.
A good positional encoding must be bounded in size, distinguish every position uniquely, and make relative offsets easy to compute. Those three requirements rule out almost every simple idea.
Sinusoidal encoding: worked out
The original transformer's answer was to give each position a fixed vector built from sine and cosine waves at geometrically spaced frequencies:
Concretely, with d=4 there are two frequency pairs. For i=0 the divisor is 100000=1; for i=1 it is 100002/4=100. So the vector for position pos is
| Position | dim 0: sin(pos) | dim 1: cos(pos) | dim 2: sin(pos/100) | dim 3: cos(pos/100) |
|---|---|---|---|---|
| 0 | 0.000 | 1.000 | 0.000 | 1.000 |
| 1 | 0.841 | 0.540 | 0.010 | 1.000 |
| 2 | 0.909 | −0.416 | 0.020 | 1.000 |
| 3 | 0.141 | −0.990 | 0.030 | 1.000 |
Two properties fall straight out of the table. Every component stays within [−1,1], so nothing explodes at position 100,000. And the columns vary at wildly different rates: dimension 0 swings across a full cycle every ~6 positions, dimension 2 needs ~628 positions to do the same. The fast dimensions carry fine-grained local position; the slow ones carry coarse "roughly where in the document" information. Together they act like the hands of a clock — several dials at different speeds, jointly pinning down a unique reading.
The deeper reason for sinusoids is a trigonometric identity. Because
the encoding of position pos+k is a fixed linear function of the encoding of position pos — a rotation by an angle that depends only on the offset k. So a linear layer, which is all attention has available, can in principle learn "look 3 back" as a single fixed transformation that works at every position. That is exactly the relative-offset property the naive index scheme lacked.
Learned absolute positions, and where they stop
GPT-2 and BERT took a blunter route: make positions another embedding table. A matrix of shape (max_length × dmodel), trained by gradient descent like everything else, and simply added to the token embedding.
1import torch, torch.nn as nn23class Embed(nn.Module):4 def __init__(self, vocab, d_model, max_len):5 super().__init__()6 self.tok = nn.Embedding(vocab, d_model)7 self.pos = nn.Embedding(max_len, d_model) # learned, one row per slot89 def forward(self, ids): # ids: (batch, seq)10 seq = ids.shape[1]11 positions = torch.arange(seq, device=ids.device)12 return self.tok(ids) + self.pos(positions) # broadcast over batchThis works well inside the trained range and costs almost nothing to implement. Its limitation is absolute: row 1,024 of a table with 1,024 rows does not exist. A model trained with a 1,024-token limit cannot be handed 1,500 tokens at all, and the rows near the end of the table — reached by comparatively few training sequences — are typically undertrained.
Rotary embeddings (RoPE), and why they took over
Both schemes above add a position vector to the token vector, mixing "what" and "where" into one signal. RoPE does something structurally different: it rotates the query and key vectors by an angle proportional to their position, leaving the value vectors alone.
Split each query and key into 2-dimensional pairs. For a token at position m, rotate the pair at frequency θ by angle mθ:
The payoff is a clean piece of geometry. Rotating one vector by mθ and another by nθ, then taking their dot product, gives a result that depends only on m−n. Watch it happen with real numbers, using θ=1 radian and both raw vectors equal to (1,0):
query at position m = 3: rotate (1,0) by 3 rad -> (cos 3, sin 3) = (-0.990, 0.141)key at position n = 1: rotate (1,0) by 1 rad -> (cos 1, sin 1) = ( 0.540, 0.841)dot product = (-0.990)(0.540) + (0.141)(0.841) = -0.5346 + 0.1186 = -0.416and cos(m - n) = cos(2) = -0.416 <- identicalMove both tokens 500 positions further into the document — m=503, n=501 — and the dot product is still cos(2)=−0.416. The attention score literally cannot see absolute position, only the gap. Relative position is not learned or approximated; it is baked into the algebra.
Three practical consequences follow. Rotation preserves vector length, so nothing blows up at long range. There is no positional table to run out of, so a longer input is always accepted — but that is not the same as working. Past its training length the model meets rotation angles, and so relative gaps, it never saw, and quality typically falls off quickly rather than gracefully. And because the position dependence lives entirely in the θ frequencies, the context window can be stretched after training by changing those frequencies — the family of tricks behind "extended context" versions of open models.
How context extension actually works
The recipes differ in which frequencies they touch, but share one shape: change the RoPE schedule so long positions map onto angles the model already knows, then continue training briefly on long documents so it adapts.
- Position interpolation divides every position by the extension factor. To go from 4k to 16k, position 12,000 is rotated as if it were position 3,000, so no angle is new. The cost is that neighbouring tokens become harder to tell apart, because every gap has shrunk by 4×.
- NTK-aware scaling and YaRN stretch the slow, low-frequency pairs (which carry "roughly where in the document") and leave the fast, high-frequency pairs (which carry "the token just before me") nearly alone, so local order stays sharp.
- A larger base. Many recent models simply train with a base far above the original 10,000 in θi=base−2i/d — Llama 3, for example, uses 500,000 — which slows every rotation down from the start.
The short continued-training step matters. Rescaling alone produces a model that accepts long inputs but uses the far end of them badly, which is the gap the advice at the end of this lesson warns about.
A fourth approach, ALiBi, is worth knowing as the minimalist option: skip positional vectors entirely and simply subtract a penalty proportional to distance from each raw attention score, −m⋅∣i−j∣ with a different slope m per head. It builds in a hard recency bias, extrapolates well, and costs essentially nothing.
| Sinusoidal | Learned absolute | RoPE | ALiBi | |
|---|---|---|---|---|
| Parameters added | 0 | max_len × d | 0 | 0 |
| Applied to | Input embedding | Input embedding | Q and K, every layer | Attention scores |
| Encodes | Absolute (relative recoverable) | Absolute only | Relative directly | Relative distance penalty |
| Beyond training length | Defined but degrades | Impossible | Accepted, but degrades quickly; extendable by rescaling plus brief long-context training | Extrapolates well |
| Typical users | Original transformer | GPT-2, BERT | Llama, Mistral, Qwen, most modern models | BLOOM, MPT |
The full input path, end to end
Putting the pieces together for a model using learned absolute positions, the journey from string to first transformer layer is:
"The cat sat" | tokenise v[464, 3797, 3332] shape: (3,) | embedding lookup v[[0.21, -0.05, ...], shape: (3, d_model) [0.87, 0.33, ...], [-0.12, 0.64, ...]] | scale by sqrt(d_model) | add positional encoding for rows 0, 1, 2 v[[0.21+PE0, ...], shape: (3, d_model) [0.87+PE1, ...], [-0.12+PE2, ...]] | vfirst transformer blockWith RoPE the diagram changes shape: nothing is added at the input, and the rotation is applied to the query and key vectors inside every attention layer instead. Which is itself a meaningful difference — additive schemes inject position once and hope it survives forty layers of mixing, while RoPE re-asserts it at every layer.
What this means when you are actually building
When a model's advertised context length gets extended — 4k to 32k, 32k to 128k — check how. If it was done by rescaling RoPE frequencies with little or no long-document training, the model will accept long inputs without crashing but its accuracy on material in the far reaches of the window is often much worse than on short inputs. Test retrieval at the actual depths you care about; do not trust the number on the box.
When you concatenate documents or build long prompts, remember that positional information is per-sequence, not per-document. Two documents packed into one sequence are, positionally, one continuous stream, and the model can and does attend across the join. If that is not what you want, separate them with a boundary token the model was trained to treat as a break.
If you are inspecting embeddings directly — clustering, nearest neighbours, similarity search — use the input embedding table, and be aware it encodes token identity, not word meaning in context. The row for bank is a single vector serving both the river and the financial sense; it is the later transformer layers that pull those apart using surrounding context. Judging a model's semantic ability from its raw embedding table alone will consistently understate it.