Course Content
LLMs Deep Dive
10 sections · 40 lessons
How are LLMs different from traditional language models?
What you need to know
The older models
An n-gram model with n = 3 sees only two words of history. It cannot connect "the account I opened in Pune last year" to a question about "it" ten words later. It also fails on any word pair it never counted — smoothing tricks help only a little.
RNNs and LSTMs replaced counting with a learned hidden state carried from word to word. Better, but the state is one fixed-size vector, information fades over long distances, and training must go word by word, so it cannot use GPUs fully.
What changed with LLMs
| Traditional LM | LLM | |
|---|---|---|
| Architecture | n-gram counts, RNN, LSTM | Transformer with self-attention |
| Size | Thousands to millions of parameters | Billions (some over a trillion) |
| Training data | Millions of words, one domain | Trillions of tokens, many domains and languages |
| History used | 2–5 words (n-gram), fading (RNN) | Thousands to about 1 million tokens |
| Units | Whole words, with an unknown-word token | Subword tokens, nothing is unknown |
| How you use it | Train one model per task | Prompt one model for many tasks |
In-context learning
The biggest practical difference. Show an LLM three examples of "complaint → category" in the prompt, and it classifies the fourth — with no weight update. An older model needed a labelled dataset and a new training run for the same result.
Parallel training
Attention looks at all positions at once, so a whole sequence is processed in one pass during training. That is what made training on trillions of tokens affordable. Recurrent models had to step through time, one token after another.
A real-life example
An e-commerce company built a search-query autocomplete in 2016 with a trigram model. Typing "red running" suggested "shoes", which was fine. But typing "gift for my dad who likes cricket and is" gave nonsense, because only the last two words, "and is", were used.
In 2026 the same company uses a small LLM for its search assistant. Given "gift for my dad who likes cricket and is 60, budget Rs 3,000", it keeps all of it in view — the relation, the interest, the age and the budget — and suggests a signed-bat replica under Rs 3,000. The team also added a new task, "extract the budget as a number", by writing one line of prompt instead of training a new model. The trade-off: each query now costs GPU time and occasionally invents a product, so results are checked against the real catalogue.
Follow-up questions to expect
- "What are emergent abilities?" — Skills such as multi-step arithmetic that appear to show up only past a certain scale. Some research argues the sudden jump is partly an artefact of how we measure, so say "appear to emerge".
- "Why did transformers replace LSTMs?" — Parallel training over the whole sequence, and a direct attention path between any two tokens instead of a fading hidden state.
- "Do people still use small or traditional models?" — Yes: for autocomplete on device, spam filters and tight latency budgets, a small model or even a classical one is cheaper and more predictable.