Course Content
LLMs Deep Dive
10 sections · 40 lessons
What is the difference between an encoder and a decoder?
What you need to know
Attention masks with 4 tokens
Take the input "block my debit card". In self-attention, each token computes a weight for every other token. The mask decides which weights are allowed.
Encoder (bidirectional) — every token sees every token:
block my debit cardblock ✓ ✓ ✓ ✓my ✓ ✓ ✓ ✓debit ✓ ✓ ✓ ✓card ✓ ✓ ✓ ✓Decoder (causal) — each token sees only itself and earlier tokens:
block my debit cardblock ✓ · · ·my ✓ ✓ · ·debit ✓ ✓ ✓ ·card ✓ ✓ ✓ ✓The blocked cells get a score of minus infinity before softmax, so their weight becomes 0. Why block? During training the decoder learns to predict "card" from "block my debit". If it could see "card", it would copy the answer and learn nothing.
Cross-attention
In an encoder–decoder model, each decoder layer also has cross-attention: its queries come from the decoder, but keys and values come from the encoder output. That is how the decoder "reads" the source sentence while writing the translation.
Which to use
| Encoder-only | Decoder-only | Encoder–decoder | |
|---|---|---|---|
| Attention | Bidirectional | Causal | Bidirectional in, causal out, plus cross |
| Examples | BERT, RoBERTa, DeBERTa, embedding models | GPT, Claude, Gemini, Llama, Qwen | T5, BART, Whisper |
| Strong at | Classification, NER, search embeddings | Generation, chat, reasoning, code | Translation, speech-to-text |
| Typical size in production | 100M–400M parameters | 1B to hundreds of billions | Tens of millions to about 10B |
Encoders are small and fast. A fine-tuned 110M-parameter BERT can classify a message in a few milliseconds on a CPU, which a large decoder cannot match.
A real-life example
An e-commerce search assistant has two jobs. First, turn each product description into a vector for semantic search: an encoder embedding model does this, because it reads the whole description bidirectionally and outputs one vector. Second, answer the shopper, "Which of these three phones has the best battery for under Rs 20,000?": a decoder LLM does this, because it must generate a new sentence.
Using the decoder LLM for embeddings would cost more and usually retrieve worse; using the encoder to write answers is impossible, since it has no way to generate text left to right.
Follow-up questions to expect
- "Can a decoder-only model produce embeddings?" — Yes; many modern embedding models are built from decoder LLMs fine-tuned with a contrastive objective, often with the causal mask adjusted. Plain hidden states from a chat model are weaker for search.
- "Why is BERT not used for chat?" — It has no causal mask and was not trained to generate, so it cannot produce fluent text token by token.
- "What is the prefix-LM mask?" — A hybrid: bidirectional attention over the input prefix, causal over the generated part. Some models such as T5 variants and UL2 used it.