Transformer Architecture Q&A

Course Content

Transformer Architecture Q&A

6 sections · 60 lessons

How does bidirectional encoder self-attention differ from causal decoder self-attention?


What you need to know

The masks side by side

For 4 tokens, "1" means "may attend":

Text
bidirectional           causal1 1 1 1                 1 0 0 01 1 1 1                 1 1 0 01 1 1 1                 1 1 1 01 1 1 1                 1 1 1 1

Everything else — projections, scaling, softmax, value mixing — is the same code.

Effect 1: representation quality

In "I sat on the bank of the Ganga", the word that decides the meaning of "bank" comes after it. A bidirectional encoder sees "Ganga" when building the vector for "bank" in layer 1. A causal model's vector for "bank" can never see "Ganga"; only later tokens can combine the two. Early tokens in a causal model have thin context — the first token sees only itself.

Effect 2: training objective

If every position can see every token, next-token prediction is trivial (the answer is in the input). So BERT-style models use masked language modelling: hide about 15% of tokens and predict them. Only those 15% give a training signal, whereas a causal model gets a signal at every position.

Effect 3: caching and generation

In a causal model, adding token 101 does not change the keys and values of tokens 1–100, because they could not see it. You compute them once and cache them. In a bidirectional model, adding one token changes every token's representation in every layer, so nothing can be cached and each new token would mean recomputing everything.

Bidirectional (encoder)

  • Every token sees the full input
  • Trained with masked language modelling
  • No KV cache; reprocess on any change
  • Best for classification, tagging, embeddings

Causal (decoder)

  • Each token sees only the past
  • Trained with next-token prediction
  • KV cache makes generation cheap
  • Best for generation and chat

A real-life example

A bank's document classifier must flag loan agreements that contain a "prepayment penalty" clause. The key sentence is often: "The borrower shall pay a fee of 2% ... if the loan is closed before 36 months." Whether "fee" is a penalty depends on words that come after it.

The team compares a 150M-parameter bidirectional encoder with mean-pooled embeddings from a similar-sized decoder. On their test set the encoder is clearly more accurate, and it runs faster because there is no generation. For the chat assistant that explains the clause to customers, they use a causal decoder model — the task there is to write text.

Follow-up questions to expect

  • "Can one model do both?" — Prefix-LM models use bidirectional attention over a prompt prefix and causal attention for the continuation. Encoder-decoder models split the two roles into separate stacks.
  • "Does a padding mask make attention causal?" — No. A padding mask only hides padding tokens; real tokens still see both directions.
  • "Why does the first token in a causal model often get a lot of attention?" — It is visible to every position, so many heads use it as a default "no-op" target, often called an attention sink.