LLMs Deep Dive

Course Content

LLMs Deep Dive

10 sections · 40 lessons

How does the Transformer architecture overcome challenges of traditional Seq2Seq models?


What you need to know

One transformer block

A transformer is a stack of identical blocks (32 to over 100 in large LLMs). Each block has two sub-layers, each wrapped in a residual connection and a normalisation layer:

  1. Self-attention — every token gathers information from other tokens.
  2. Feed-forward network (MLP) — every token is transformed independently; this is where much of the model's stored knowledge lives.
  3. Residual add — each sub-layer's output is added to its input, so information and gradients have a direct path through the stack.

Problem 1: the bottleneck

An RNN encoder turns a 5,000-token contract into, say, 512 numbers. A transformer keeps 5,000 separate vectors, one per token, and attention can read any of them at any layer. Nothing has to be squeezed.

Problem 2: sequential training

An RNN must compute step 1 before step 2 before step 3. For a 1,000-token sequence that is 1,000 dependent steps, and a GPU with thousands of cores sits mostly idle. A transformer computes all 1,000 positions at once as large matrix multiplications — exactly what GPUs and TPUs are built for. This speed-up is the real reason large-scale pretraining happened.

Note that generation is still sequential: at inference the model produces one token at a time. The parallelism is in training and in reading the prompt.

Problem 3: long-range dependencies

In an RNN, information from token 1 reaches token 500 only after passing through 499 updates, and gradients shrink on the way back. In a transformer the path length between any two tokens is one attention step, whatever the distance.

The new costs

  • Quadratic attention — each token scores against every other, so n tokens give n × n scores. 4,000 tokens is 16 million scores per head per layer; 128,000 tokens is about 16 billion — 1,024 times more for 32 times the length.
  • No built-in order — attention treats input like a set, so positional encodings are needed.
  • Memory at inference — keys and values for past tokens are cached (the KV cache), which grows with every token.

FlashAttention reduces memory traffic but keeps the quadratic compute. Some 2025–26 models mix attention layers with cheaper linear-attention or state-space layers to reduce the cost on very long inputs.

A real-life example

A legal-document summariser built on an LSTM in 2018 took about 3 days to train on 1 million contracts, because each contract was processed word by word. Summaries of long contracts lost the details from later clauses.

The team moves to a transformer. Training on the same data takes a fraction of the time because every token of a batch is processed in parallel. Later clauses are no longer forgotten — the "governing law" clause on page 28 is one attention hop from the summary it feeds. The new problem is cost: doubling the contract length from 16,000 to 32,000 tokens makes the attention part four times more expensive, so the team splits very long agreements by chapter and summarises in two stages.

Follow-up questions to expect

  • "Is generation in a transformer parallel?" — No. Reading the prompt (prefill) is parallel, but output tokens are produced one at a time, each needing a full forward pass.
  • "What does the feed-forward layer do?" — It applies the same two-layer network to each token independently, adding capacity; research suggests many factual associations are stored there.
  • "Why are residual connections important?" — They give an identity path, so deep stacks train stably and each layer only needs to learn a correction.