Course Content
Transformer Architecture Q&A
6 sections · 60 lessons
Identify where dropout is typically applied in GPT architecture?
What you need to know
The three places
In Hugging Face's GPT-2 config these are embd_pdrop, attn_pdrop and resid_pdrop:
x = dropout_embd(tok_emb + pos_emb) # 1. embedding dropoutfor each block: a = softmax(QKᵀ/sqrt(d) + mask) a = dropout_attn(a) # 2. attention dropout x = x + dropout_resid(W_o(a · V)) # 3a. residual dropout x = x + dropout_resid(MLP(LN2(x))) # 3b. residual dropoutlogits = LN_f(x) · Wᵀ # no dropout here(Layer norm before attention, LN1, is left out of the sketch for space.)
Why exactly these places
- Embedding dropout — stops the model relying on any one dimension of a token's embedding.
- Attention dropout — stops a head relying on one source token.
- Residual dropout — drops parts of each sublayer's update, so no single block becomes a critical path. The residual stream itself stays intact, so the "identity path" through the network is always available.
Where it is never applied
- The residual stream itself — dropping it would randomly delete information that every later layer needs, and break the shortcut that makes deep stacks trainable.
- The logits — dropping them would randomly remove candidate tokens from the loss.
- At inference —
model.eval()turnsnn.Dropoutinto the identity.
Current practice
| Model family | Dropout rates |
|---|---|
| GPT-2 (2019) | 0.1 at all three places |
| Llama 2 / 3, Mistral, Qwen | 0.0 in their configs (for example attention_dropout: 0.0) |
| Fine-tuning small data | re-enable around 0.1, or use LoRA dropout |
Dropout also costs speed: it needs random masks and extra memory traffic, and at the scale of trillion-token pretraining that is a real cost for little benefit.
A real-life example
An Indian-language news app fine-tunes a small GPT-2-sized model to write Hindi headlines from story summaries, using 20,000 pairs. With dropout off (the checkpoint's setting), training loss keeps falling but the validation loss rises after the second epoch, and headlines start to repeat phrases from the training set word for word.
They set embd_pdrop, attn_pdrop and resid_pdrop to 0.1. Training loss now falls more slowly, validation loss keeps improving into the fourth epoch, and editors rate more headlines as usable. For a later 8B model they switch to LoRA: the base model's dropout stays at 0, and LoRA's own dropout of 0.05 is applied only on the adapter path.
Follow-up questions to expect
- "Why not dropout inside the MLP between the two linear layers?" — Some implementations do; GPT-2 applies it at the MLP output instead, which regularises the whole sublayer's update.
- "Why do large LLMs not use dropout?" — Each token is seen about once in pretraining, so memorisation is limited; dropout would slow training and lower effective capacity.
- "Does dropout make generation random?" — No. It is disabled at inference; randomness in generation comes only from sampling.