Course Content
Transformer Architecture Q&A
6 sections · 60 lessons
What does "autoregressive" mean in Transformers, and how does it differ in GPT vs BERT?
What you need to know
p(x_1, x_2, ..., x_T) = p(x_1) · p(x_2 | x_1) · p(x_3 | x_1, x_2) · ... · p(x_T | x_1 .. x_{T-1})Worked by hand
Suppose a model assigns these probabilities for "chai garam hai":
p("chai") = 0.2p("garam" | "chai") = 0.5p("hai" | "chai garam") = 0.6p(sentence) = 0.2 × 0.5 × 0.6 = 0.06Training maximises the log of this product, which is the sum of the three log-probabilities — so every token is a training target. Generation samples x_1, then x_2 given x_1, and so on.
GPT: autoregressive by design
The causal mask enforces the chain: position 2 can only see position 1. One forward pass yields all the conditionals for training. At inference each new token needs one more forward pass, which is why decoding speed is measured in tokens per second.
BERT: a masked denoiser
BERT picks 15% of tokens; of those, 80% become [MASK], 10% become a random token and 10% stay unchanged. It predicts the originals using both sides:
input : chai [MASK] haitarget: predict "garam" using "chai" AND "hai"This models p(masked tokens | visible tokens), not a left-to-right chain. You cannot sample text from it in order, because each prediction assumes the right-hand context is already known. You can hack generation by repeated masking and filling, but quality is poor.
A 2026 note
Research on masked diffusion language models revisits BERT-like denoising for generation: start from all masks and unmask many tokens over several steps. It is an active area, but mainstream production LLMs remain autoregressive.
A real-life example
A speech-to-text system for a Hindi news channel's archive uses Whisper. Its decoder is autoregressive: after the encoder processes 30 seconds of audio, the decoder writes the transcript one token at a time — "आज", then "दिल्ली", then "में" — each conditioned on the audio and all earlier text tokens.
Two practical results follow. First, a 200-token transcript needs 200 decoder steps, so decoder speed dominates latency; the team uses a smaller distilled decoder to go faster. Second, one mis-heard word early in a segment changes the context for every later word, so errors can cascade — beam search, which keeps several candidate sequences, reduces that. A BERT-style model could not do this job at all: there is no text yet to mask.
Follow-up questions to expect
- "Is teacher forcing autoregressive?" — The factorisation is; training just computes all conditionals in parallel using the true previous tokens instead of the model's own.
- "What is exposure bias?" — In training the model always sees correct history; at inference it sees its own mistakes, which it never practised recovering from.
- "Is an encoder-decoder autoregressive?" — Its decoder is. The encoder is not; it reads the input all at once.