LLMs Deep Dive

Course Content

LLMs Deep Dive

10 sections · 40 lessons

What is the difference between an encoder and a decoder?


The causal mask a decoder applies✓···✓✓··✓✓✓·✓✓✓✓blockmydebitcardblockmydebitcardAn encoder allows every cell; the dotted cells are set to minus infinity before softmax.
Blocking the upper triangle is what stops a decoder from copying the word it is being trained to predict.

What you need to know

Attention masks with 4 tokens

Take the input "block my debit card". In self-attention, each token computes a weight for every other token. The mask decides which weights are allowed.

Encoder (bidirectional) — every token sees every token:

Text
            block  my  debit  cardblock         ✓    ✓     ✓     ✓my            ✓    ✓     ✓     ✓debit         ✓    ✓     ✓     ✓card          ✓    ✓     ✓     ✓

Decoder (causal) — each token sees only itself and earlier tokens:

Text
            block  my  debit  cardblock         ✓    ·     ·     ·my            ✓    ✓     ·     ·debit         ✓    ✓     ✓     ·card          ✓    ✓     ✓     ✓

The blocked cells get a score of minus infinity before softmax, so their weight becomes 0. Why block? During training the decoder learns to predict "card" from "block my debit". If it could see "card", it would copy the answer and learn nothing.

Cross-attention

In an encoder–decoder model, each decoder layer also has cross-attention: its queries come from the decoder, but keys and values come from the encoder output. That is how the decoder "reads" the source sentence while writing the translation.

Which to use

Encoder-onlyDecoder-onlyEncoder–decoder
AttentionBidirectionalCausalBidirectional in, causal out, plus cross
ExamplesBERT, RoBERTa, DeBERTa, embedding modelsGPT, Claude, Gemini, Llama, QwenT5, BART, Whisper
Strong atClassification, NER, search embeddingsGeneration, chat, reasoning, codeTranslation, speech-to-text
Typical size in production100M–400M parameters1B to hundreds of billionsTens of millions to about 10B

Encoders are small and fast. A fine-tuned 110M-parameter BERT can classify a message in a few milliseconds on a CPU, which a large decoder cannot match.

A real-life example

An e-commerce search assistant has two jobs. First, turn each product description into a vector for semantic search: an encoder embedding model does this, because it reads the whole description bidirectionally and outputs one vector. Second, answer the shopper, "Which of these three phones has the best battery for under Rs 20,000?": a decoder LLM does this, because it must generate a new sentence.

Using the decoder LLM for embeddings would cost more and usually retrieve worse; using the encoder to write answers is impossible, since it has no way to generate text left to right.

Follow-up questions to expect

  • "Can a decoder-only model produce embeddings?" — Yes; many modern embedding models are built from decoder LLMs fine-tuned with a contrastive objective, often with the causal mask adjusted. Plain hidden states from a chat model are weaker for search.
  • "Why is BERT not used for chat?" — It has no causal mask and was not trained to generate, so it cannot produce fluent text token by token.
  • "What is the prefix-LM mask?" — A hybrid: bidirectional attention over the input prefix, causal over the generated part. Some models such as T5 variants and UL2 used it.