Course Content
Transformer Architecture Q&A
6 sections · 60 lessons
Compare encoder-only, decoder-only, encoder-decoder Transformers and their best use cases?
What you need to know
| Encoder-only | Decoder-only | Encoder-decoder | |
|---|---|---|---|
| Attention | Bidirectional | Causal | Bidirectional encoder, causal decoder, plus cross-attention |
| Training objective | Masked language modelling | Next-token prediction | Denoising or sequence-to-sequence |
| Output | A vector per token | Next-token distribution | A generated target sequence |
| Examples | BERT, RoBERTa, ModernBERT | GPT, Llama, Mistral, Qwen | Original Transformer, T5, Whisper |
Encoder-only
Each token sees the whole input, left and right, in every layer. That gives rich representations for understanding tasks. You add a small head on top: a linear layer over a pooled vector for classification, or one per token for tagging. Encoders are small and fast: a base-size encoder of about 150M parameters classifies a short document in a few milliseconds on one GPU. They cannot generate fluent text.
Decoder-only
Each token sees only earlier tokens. One objective — predict the next token — scales to trillions of tokens, and in-context learning lets one model handle tasks that once needed separate models. The KV cache makes generation efficient. This is the design of nearly all 2026 chat and reasoning LLMs.
Encoder-decoder
The encoder reads the input once, bidirectionally. The decoder generates the output, attending to its own past tokens (causal) and to the encoder output (cross-attention). This fits tasks where the input is complete and the output is a transformation of it. The encoder output is computed once and reused for every output token, which is efficient when the input is long and the output short. The cost is two stacks to train and serve, and a more complex inference path.
When the answer changes
Large decoder-only LLMs can now do translation and classification well through prompting. The dedicated architectures win when cost, latency or a narrow task matters more than generality.
A real-life example
One company runs four systems and uses a different variant for each:
- Document classifier that routes 2 million insurance PDFs a month into claim types: a fine-tuned encoder (ModernBERT-size). It runs in milliseconds and needs no generation.
- Code-completion assistant in the IDE: a decoder-only model, because it must write code token by token and reuse the KV cache as the user types.
- Indian-language news app translating English stories into Tamil and Bengali: an encoder-decoder translation model such as IndicTrans2 for bulk translation, with a general LLM as a fallback for headlines that need a particular style.
- Speech-to-text for call-centre audio: Whisper, an encoder-decoder. The encoder reads 30 seconds of audio; the decoder writes the transcript while cross-attending to the audio.
Choosing a large decoder LLM for the classifier would work, but at many times the cost per document.
Follow-up questions to expect
- "Why did decoder-only win for general LLMs?" — One simple objective uses every token as a training signal, scales cleanly, and handles many tasks by prompting.
- "Are encoders obsolete?" — No. Embedding models, rerankers and fast classifiers are still mostly encoders, because bidirectional context is cheap and effective for understanding.
- "Can a decoder produce embeddings?" — Yes, and many modern embedding models are adapted decoders, but they need extra training because causal attention gives early tokens little context.