Transformer Architecture Q&A

Course Content

Transformer Architecture Q&A

6 sections · 60 lessons

What does GPT stand for, and what is the architectural intent behind it?


What you need to know

Each word is a design decision

  • Generative — the model learns p(next token | all previous tokens). It can therefore write text, one token at a time. A BERT-style encoder, by contrast, outputs a vector per token and must have a task head added on top.
  • Pre-trained — the objective needs no labels: every position in every document is a training example ("predict the next token"). The labelled or preference-based work comes later.
  • Transformer (decoder-only) — a stack of identical blocks, each with causal self-attention and an FFN, each writing into a residual stream. No encoder, so there is nothing to connect with cross-attention.

How the idea developed

ModelYearSizeWhat it showed
GPT-1 (Radford et al.)2018~117M, 12 layersPre-train on books, then fine-tune per task, beats task-specific models
GPT-22019up to 1.5BWith enough scale, many tasks work zero-shot from a prompt
GPT-32020175BFew-shot learning in the prompt; no fine-tuning needed
InstructGPT / ChatGPT2022—Fine-tuning on instructions and human preferences (RLHF) makes it a helpful assistant

In 2026, "pre-trained" is only the first stage. Modern GPT-style models add supervised fine-tuning, preference tuning (RLHF or DPO), and often reinforcement learning on checkable tasks for reasoning. The architecture stays a decoder-only Transformer.

Why "decoder-only" won

  • One objective uses every token. A 1,000-token document gives 1,000 training predictions in one pass.
  • One homogeneous stack. Scaling is "more of the same block", which suits the engineering of very large training runs.
  • One interface. Translation, summarisation, code and chat all become "continue this text". The same model and serving stack handle all of them.

Encoder-only models (BERT-style) still win for tasks that only need a representation — search, classification — because they read the whole input in both directions at lower cost.

A real-life example

An Indian-language news app once ran three separate systems: an encoder-decoder model for English-to-Hindi translation, a BERT classifier for tagging, and a template system for headlines. In 2026 the team replaces the first and third with one decoder-only model: the prompt says "Translate this story into Hindi, then write a 12-word headline". One model, one serving stack, one set of evaluations.

They keep the BERT-style classifier for tagging, because it runs on CPU in a few milliseconds per article and a generative model would be slower and costlier for a job that needs no text output. That split — decoders for generation, encoders for cheap understanding — is the practical meaning of "generative".

Follow-up questions to expect

  • "How is GPT different from BERT?" — GPT is a causal decoder trained to predict the next token and can generate; BERT is a bidirectional encoder trained to fill in masked tokens and produces representations.
  • "Is ChatGPT just a pre-trained GPT?" — No. It is a pre-trained GPT plus instruction tuning and preference tuning, which change its behaviour, not its architecture.
  • "Why no encoder?" — For a model that only continues text, the prompt can be read by the same causal stack, so a separate encoder and cross-attention add parameters without adding capability.