Course Content
Transformer Architecture Q&A
6 sections · 60 lessons
List the major building blocks of a Transformer and briefly state each role?
What you need to know
Tokenizer
Models do not read characters or whole words. A tokenizer, usually byte-pair encoding (BPE), learns frequent sub-word pieces from data. A common word is one token; a rare word is split into pieces, for example unbelievable might become un + believ + able. Each piece has an integer id. The tokenizer is not trained with the network, but it is fixed for the model's life.
Embedding table
The embedding step is a lookup, not a calculation: id 2 selects row 2 of a learned matrix E of shape (V, d_model).
1import numpy as np23vocab = {"<unk>": 0, "chai": 1, "garam": 2, "hai": 3, "thandi": 4}4d_model = 45rng = np.random.default_rng(0)6E = rng.normal(0, 0.02, size=(len(vocab), d_model)) # (V, d_model) learned table78ids = [vocab[w] for w in "chai garam hai".split()] # [1, 2, 3]9x = E[ids] # row lookup, no maths10print(ids, x.shape) # [1, 2, 3] (3, 4)During training, gradients update only the rows that were used. In GPT-2 small the table is 50,257 × 768, about 38.6 million parameters.
The rest of the stack
| Block | Role | What breaks without it |
|---|---|---|
| Positional information | Tells the model token order | "India beat Australia" equals "Australia beat India" |
| Multi-head self-attention | Moves information between tokens | Each token is processed alone |
| Causal mask (decoders) | Hides future tokens | Training cheats by reading the answer |
| Feed-forward network | Per-token non-linear processing, stores much knowledge | The model is nearly linear and weak |
| Residual connections | Adds each sub-layer's output to its input | Deep stacks do not train |
| Normalisation (LayerNorm or RMSNorm) | Keeps vector scale stable | Activations drift, training diverges |
| Output head | Maps d_model to V logits | No prediction |
Modern choices: RoPE for position, RMSNorm placed before each sub-layer (pre-norm), SwiGLU in the FFN, and grouped-query attention (GQA) to shrink the KV cache.
A real-life example
A company builds two systems. One is a document classifier that sorts incoming PDFs into invoice, contract or ID proof. It uses an encoder model (a BERT-style model): tokenizer, embeddings, positions, bidirectional attention, FFN, residuals and norms — but no causal mask and no LM head. Instead a small classification head reads one pooled vector and outputs three scores.
The other is a code-completion assistant built on a decoder-only model. It uses all the same blocks plus the causal mask and the LM head. Listing blocks this way shows the interviewer you know which pieces are shared and which are task-specific.
Follow-up questions to expect
- "What is weight tying?" — Using the embedding matrix (transposed) as the output head, which saves
V × d_modelparameters. GPT-2 does this; many large modern models keep them separate. - "Why sub-words and not whole words?" — A word vocabulary cannot cover new names or typos; characters make sequences very long. Sub-words balance both.
- "Which block holds the most parameters?" — The FFNs, typically about two-thirds or more of non-embedding weights.