LLMs Deep Dive

Course Content

LLMs Deep Dive

10 sections · 40 lessons

What are Sequence-to-Sequence (Seq2Seq) Models?


What you need to know

The shape of the problem

Many tasks are "sequence in, different sequence out":

  • Translation: 12 Hindi words in, 10 English words out.
  • Summarisation: 8,000 tokens in, 300 out.
  • Speech recognition: audio frames in, text out.

A normal classifier outputs one label. A seq2seq model must output a sequence whose length it decides itself, stopping when it produces an end-of-sequence token.

The original RNN design (2014)

  1. Encode — an RNN reads the input token by token, updating a hidden state.
  2. Compress — the final hidden state (say 512 numbers) becomes the "context vector" for the whole input.
  3. Decode — a second RNN starts from that vector and emits output tokens one at a time, feeding each back in.

The flaw is step 2. A 5-word sentence and a 5,000-token contract both have to fit in the same 512 numbers. Quality dropped sharply as inputs got longer.

Attention fixes the bottleneck

In 2015, attention let the decoder look back at every encoder hidden state at each step and take a weighted mix of the relevant ones. When writing the English word "account", the decoder can focus on the Hindi word "खाता". The transformer (2017) kept attention and removed recurrence completely, which also made training parallel.

Where seq2seq stands in 2026

DesignExamplesTypical use
Encoder–decoder transformerT5, BART, WhisperTranslation, speech-to-text, some summarisation
Decoder-only transformerGPT, Claude, Gemini, Llama, QwenEverything, including seq2seq tasks via a prompt

A decoder-only model does seq2seq by putting the input in the prompt ("Translate to English: ...") and generating the output after it. With enough scale this works as well or better, and one architecture serves all tasks, which is why it dominates.

A real-life example

A legal-tech startup built a contract summariser in 2017 with an RNN seq2seq model. On 1-page NDAs it was fine. On 30-page supply agreements, summaries mentioned the parties from page 1 and then drifted into generic text, because everything after the first few pages had been squeezed out of the context vector.

The 2026 version sends the whole agreement (about 20,000 tokens) to a decoder-only LLM with a long context window. When writing "Termination requires 90 days' notice", attention reaches directly back to clause 14.2. The team still checks each claim in the summary against the cited clause, because the model can summarise a clause fluently and wrongly.

Follow-up questions to expect

  • "What is teacher forcing?" — During training, the decoder is fed the correct previous token rather than its own prediction, so all positions train in parallel. At inference it must use its own outputs, which can compound errors.
  • "How does the decoder know when to stop?" — It learns to emit an end-of-sequence token; a maximum-length limit acts as a safety cap.
  • "Why are most LLMs decoder-only instead of encoder–decoder?" — One stack is simpler to scale, next-token prediction uses every token as a training signal, and prompting turns any task into continuation.