LLMs Deep Dive

Course Content

LLMs Deep Dive

10 sections · 40 lessons

How do autoregressive models differ from masked models?


What you need to know

Autoregressive objective

For the sentence "your card is now blocked", training creates one prediction per position:

Text
"your"                    -> predict "card""your card"               -> predict "is""your card is"            -> predict "now""your card is now"        -> predict "blocked"

All four predictions are computed in one pass using the causal mask, so every token in the corpus is a training example. At inference, the model generates the same way: predict, append, repeat.

Masked objective

Text
Input:  "your [MASK] is now blocked"Target: predict "card" at position 2, using "your" AND "is now blocked"

Seeing both sides gives better representations for understanding. But only the 15% of masked tokens produce a training signal, and there is no natural way to write a paragraph one token at a time.

Other objectives worth knowing

  • Span corruption (T5) — hide whole spans and generate them; a middle ground for encoder–decoder models.
  • Denoising (BART) — shuffle, delete or mask text, then reconstruct the original.
  • Fill-in-the-middle — code models are trained to fill a gap given the code before and after it, which is how editor autocomplete works mid-file.
  • Diffusion language models — an experimental line that starts from fully masked text and unmasks many tokens in parallel over several steps. They are not yet the mainstream choice.

Pretraining is only the first stage

Both kinds of model are then adapted: masked encoders are fine-tuned with a small classification head; autoregressive models go through instruction tuning and preference tuning (RLHF or DPO) to become chat assistants.

A real-life example

A bank receives 200,000 support messages a day. Step one is routing into 12 queues (card block, KYC, loan, fraud...). The team fine-tunes a 110M-parameter masked encoder on 20,000 labelled messages. It runs on CPU in about 5 ms per message and reaches the accuracy they need.

Step two is drafting a reply for the agent. That needs generation, so an autoregressive LLM writes the draft using the customer's message and account facts. Using the LLM for routing too would work, but at 200,000 calls a day it would cost far more and add hundreds of milliseconds per message for no accuracy gain.

Follow-up questions to expect

  • "Why did the industry move to autoregressive models?" — Every token is a training signal, generation is natural, and prompting lets one model do all tasks, so scaling paid off more.
  • "What is XLNet?" — A 2019 model that trained on random orderings of the sequence (permutation language modelling) to get bidirectional context without [MASK] tokens. Historically interesting, rarely used now.
  • "Is BERT obsolete?" — Not for production classification and embeddings, where small encoders (and their modern successors such as ModernBERT) are cheap and accurate.