Course Content
Deep Learning Essentials
13 sections · 61 lessons
What are Seq2Seq (encoder-decoder) models? How is it different from Autoencoders?
What you need to know
How a seq2seq model runs
- Encode — the encoder reads all input tokens ("Where is the station?") and produces one vector per token.
- Decode, step 1 — the decoder starts with a start token and predicts the first output token.
- Decode, step 2 onwards — it feeds back what it has produced and predicts the next token, until it outputs an end token.
During training, the decoder is given the correct previous tokens instead of its own guesses. This is teacher forcing, and it makes training much faster and more stable.
The bottleneck problem, and attention
The first seq2seq models (2014) used LSTMs and passed only the encoder's final hidden state to the decoder. A 50-word sentence had to fit into one vector of, say, 512 numbers, and quality dropped for long sentences.
Attention fixed this: at every output step, the decoder computes a weighted average over all encoder vectors, focusing on the input words that matter for the current output word. Transformers (2017) built the whole model from attention. Here are the shapes in PyTorch:
1import torch, torch.nn as nn23model = nn.Transformer(d_model=128, nhead=4, num_encoder_layers=2,4 num_decoder_layers=2, batch_first=True)5src = torch.randn(8, 20, 128) # 8 English sentences, 20 tokens each (already embedded)6tgt = torch.randn(8, 14, 128) # the Hindi output so far, 14 tokens7mask = nn.Transformer.generate_square_subsequent_mask(14) # no peeking ahead8print(model(src, tgt, tgt_mask=mask).shape) # torch.Size([8, 14, 128])The input has 20 positions and the output 14: the lengths are independent. The mask stops each output position from seeing future output tokens.
Seq2seq vs autoencoder
Seq2seq model
- Target is a different sequence
- Input and output lengths can differ
- Goal: transform (translate, summarise)
- Needs paired examples
Autoencoder
- Target is the input itself
- Output has the same shape as input
- Goal: compress and represent
- Needs only unlabelled data
Both are encoder-decoder architectures. The objective is what separates them.
A real-life example
A railway booking chatbot must handle Hinglish messages like "kal Delhi se Jaipur ki train chahiye 2 log". The team trains a small seq2seq transformer to rewrite each message into a structured query: from=Delhi to=Jaipur date=tomorrow passengers=2.
The input and output have different lengths and different vocabularies. That is seq2seq. They train on 60,000 message and query pairs from past bookings, with teacher forcing.
Separately, the same team uses an autoencoder on the message embeddings to flag unusual messages (spam, abuse, or queries in a new language) by their reconstruction error. Same encoder-decoder shape, completely different purpose.
Follow-up questions to expect
- "What is teacher forcing, and its downside?" — Feeding the true previous token during training. At inference the model sees its own, possibly wrong, tokens, so errors can compound; this mismatch is called exposure bias.
- "Are GPT models seq2seq?" — GPT is decoder-only: it treats input and output as one sequence. T5 and the original transformer are encoder-decoder. Both can do seq2seq tasks.
- "How is the output generated at inference?" — Token by token with greedy decoding, beam search, or sampling, until an end token or length limit.