Transformer Architecture Q&A

Course Content

Transformer Architecture Q&A

6 sections · 60 lessons

Identify where dropout is typically applied in GPT architecture?


Where GPT-2 drops, and where it never doesEmbedding sum: embd_pdrop 0.1Post-softmax attention weights: attn_pdropSublayer outputs before the add: resid_pdropNever: residual stream, logits, inference
Dropout touches each update to the residual stream but never the stream itself, so the identity path through the network always survives.

What you need to know

The three places

In Hugging Face's GPT-2 config these are embd_pdrop, attn_pdrop and resid_pdrop:

Text
x = dropout_embd(tok_emb + pos_emb)              # 1. embedding dropoutfor each block:    a = softmax(QKᵀ/sqrt(d) + mask)    a = dropout_attn(a)                           # 2. attention dropout    x = x + dropout_resid(W_o(a · V))            # 3a. residual dropout    x = x + dropout_resid(MLP(LN2(x)))           # 3b. residual dropoutlogits = LN_f(x) · Wᵀ                              # no dropout here

(Layer norm before attention, LN1, is left out of the sketch for space.)

Why exactly these places

  • Embedding dropout — stops the model relying on any one dimension of a token's embedding.
  • Attention dropout — stops a head relying on one source token.
  • Residual dropout — drops parts of each sublayer's update, so no single block becomes a critical path. The residual stream itself stays intact, so the "identity path" through the network is always available.

Where it is never applied

  • The residual stream itself — dropping it would randomly delete information that every later layer needs, and break the shortcut that makes deep stacks trainable.
  • The logits — dropping them would randomly remove candidate tokens from the loss.
  • At inference — model.eval() turns nn.Dropout into the identity.

Current practice

Model familyDropout rates
GPT-2 (2019)0.1 at all three places
Llama 2 / 3, Mistral, Qwen0.0 in their configs (for example attention_dropout: 0.0)
Fine-tuning small datare-enable around 0.1, or use LoRA dropout

Dropout also costs speed: it needs random masks and extra memory traffic, and at the scale of trillion-token pretraining that is a real cost for little benefit.

A real-life example

An Indian-language news app fine-tunes a small GPT-2-sized model to write Hindi headlines from story summaries, using 20,000 pairs. With dropout off (the checkpoint's setting), training loss keeps falling but the validation loss rises after the second epoch, and headlines start to repeat phrases from the training set word for word.

They set embd_pdrop, attn_pdrop and resid_pdrop to 0.1. Training loss now falls more slowly, validation loss keeps improving into the fourth epoch, and editors rate more headlines as usable. For a later 8B model they switch to LoRA: the base model's dropout stays at 0, and LoRA's own dropout of 0.05 is applied only on the adapter path.

Follow-up questions to expect

  • "Why not dropout inside the MLP between the two linear layers?" — Some implementations do; GPT-2 applies it at the MLP output instead, which regularises the whole sublayer's update.
  • "Why do large LLMs not use dropout?" — Each token is seen about once in pretraining, so memorisation is limited; dropout would slow training and lower effective capacity.
  • "Does dropout make generation random?" — No. It is disabled at inference; randomness in generation comes only from sampling.