Course Content
Transformer Architecture Q&A
6 sections · 60 lessons
What does GPT stand for, and what is the architectural intent behind it?
What you need to know
Each word is a design decision
- Generative — the model learns
p(next token | all previous tokens). It can therefore write text, one token at a time. A BERT-style encoder, by contrast, outputs a vector per token and must have a task head added on top. - Pre-trained — the objective needs no labels: every position in every document is a training example ("predict the next token"). The labelled or preference-based work comes later.
- Transformer (decoder-only) — a stack of identical blocks, each with causal self-attention and an FFN, each writing into a residual stream. No encoder, so there is nothing to connect with cross-attention.
How the idea developed
| Model | Year | Size | What it showed |
|---|---|---|---|
| GPT-1 (Radford et al.) | 2018 | ~117M, 12 layers | Pre-train on books, then fine-tune per task, beats task-specific models |
| GPT-2 | 2019 | up to 1.5B | With enough scale, many tasks work zero-shot from a prompt |
| GPT-3 | 2020 | 175B | Few-shot learning in the prompt; no fine-tuning needed |
| InstructGPT / ChatGPT | 2022 | — | Fine-tuning on instructions and human preferences (RLHF) makes it a helpful assistant |
In 2026, "pre-trained" is only the first stage. Modern GPT-style models add supervised fine-tuning, preference tuning (RLHF or DPO), and often reinforcement learning on checkable tasks for reasoning. The architecture stays a decoder-only Transformer.
Why "decoder-only" won
- One objective uses every token. A 1,000-token document gives 1,000 training predictions in one pass.
- One homogeneous stack. Scaling is "more of the same block", which suits the engineering of very large training runs.
- One interface. Translation, summarisation, code and chat all become "continue this text". The same model and serving stack handle all of them.
Encoder-only models (BERT-style) still win for tasks that only need a representation — search, classification — because they read the whole input in both directions at lower cost.
A real-life example
An Indian-language news app once ran three separate systems: an encoder-decoder model for English-to-Hindi translation, a BERT classifier for tagging, and a template system for headlines. In 2026 the team replaces the first and third with one decoder-only model: the prompt says "Translate this story into Hindi, then write a 12-word headline". One model, one serving stack, one set of evaluations.
They keep the BERT-style classifier for tagging, because it runs on CPU in a few milliseconds per article and a generative model would be slower and costlier for a job that needs no text output. That split — decoders for generation, encoders for cheap understanding — is the practical meaning of "generative".
Follow-up questions to expect
- "How is GPT different from BERT?" — GPT is a causal decoder trained to predict the next token and can generate; BERT is a bidirectional encoder trained to fill in masked tokens and produces representations.
- "Is ChatGPT just a pre-trained GPT?" — No. It is a pre-trained GPT plus instruction tuning and preference tuning, which change its behaviour, not its architecture.
- "Why no encoder?" — For a model that only continues text, the prompt can be read by the same causal stack, so a separate encoder and cross-attention add parameters without adding capability.