Transformer Architecture Q&A

Course Content

Transformer Architecture Q&A

6 sections · 60 lessons

List the major building blocks of a Transformer and briefly state each role?


What you need to know

Tokenizer

Models do not read characters or whole words. A tokenizer, usually byte-pair encoding (BPE), learns frequent sub-word pieces from data. A common word is one token; a rare word is split into pieces, for example unbelievable might become un + believ + able. Each piece has an integer id. The tokenizer is not trained with the network, but it is fixed for the model's life.

Embedding table

The embedding step is a lookup, not a calculation: id 2 selects row 2 of a learned matrix E of shape (V, d_model).

Python
import numpy as npvocab = {"<unk>": 0, "chai": 1, "garam": 2, "hai": 3, "thandi": 4}d_model = 4rng = np.random.default_rng(0)E = rng.normal(0, 0.02, size=(len(vocab), d_model))   # (V, d_model) learned tableids = [vocab[w] for w in "chai garam hai".split()]    # [1, 2, 3]x = E[ids]                                            # row lookup, no mathsprint(ids, x.shape)                                   # [1, 2, 3] (3, 4)

During training, gradients update only the rows that were used. In GPT-2 small the table is 50,257 × 768, about 38.6 million parameters.

The rest of the stack

BlockRoleWhat breaks without it
Positional informationTells the model token order"India beat Australia" equals "Australia beat India"
Multi-head self-attentionMoves information between tokensEach token is processed alone
Causal mask (decoders)Hides future tokensTraining cheats by reading the answer
Feed-forward networkPer-token non-linear processing, stores much knowledgeThe model is nearly linear and weak
Residual connectionsAdds each sub-layer's output to its inputDeep stacks do not train
Normalisation (LayerNorm or RMSNorm)Keeps vector scale stableActivations drift, training diverges
Output headMaps d_model to V logitsNo prediction

Modern choices: RoPE for position, RMSNorm placed before each sub-layer (pre-norm), SwiGLU in the FFN, and grouped-query attention (GQA) to shrink the KV cache.

A real-life example

A company builds two systems. One is a document classifier that sorts incoming PDFs into invoice, contract or ID proof. It uses an encoder model (a BERT-style model): tokenizer, embeddings, positions, bidirectional attention, FFN, residuals and norms — but no causal mask and no LM head. Instead a small classification head reads one pooled vector and outputs three scores.

The other is a code-completion assistant built on a decoder-only model. It uses all the same blocks plus the causal mask and the LM head. Listing blocks this way shows the interviewer you know which pieces are shared and which are task-specific.

Follow-up questions to expect

  • "What is weight tying?" — Using the embedding matrix (transposed) as the output head, which saves V × d_model parameters. GPT-2 does this; many large modern models keep them separate.
  • "Why sub-words and not whole words?" — A word vocabulary cannot cover new names or typos; characters make sequences very long. Sub-words balance both.
  • "Which block holds the most parameters?" — The FFNs, typically about two-thirds or more of non-embedding weights.