LLMs Deep Dive

Course Content

LLMs Deep Dive

10 sections · 40 lessons

What are Large Language Models (LLMs)?


From text predictor to chat assistantPretraining — predict the next tokenInstruction tuning — follow requestsPreference tuning — RLHF or DPOReasoning RL — reward checkable answers
Every stage after pretraining shapes behaviour, not fresh knowledge — which is why the bank's live rate table still has to come from the prompt.

What you need to know

Tokens and next-token prediction

An LLM does not read words. It reads tokens — pieces of words, usually 3–4 characters of English each. At every step it outputs a probability for every token in its vocabulary (often 100,000 to 200,000 of them) and one is picked. That token is appended to the input, and the model runs again. A 300-token answer is 300 runs of the model.

Text
Input:  "Your UPI payment of Rs 500 has been"Output: " successfully" 0.61 | " debited" 0.22 | " declined" 0.09 | ...

Why "large" matters

Predicting the next token well on internet-scale text forces the model to learn a lot: that "Rs 500" is money, that a question expects an answer, that Python needs indentation. Small models learn surface patterns. Larger models trained on more data learn more general skills, which is why one model can do tasks it was never explicitly trained for.

How a chat model is made

  1. Pretraining — next-token prediction on trillions of tokens of web text, books and code. Produces a base model that continues text but does not follow instructions well.
  2. Instruction tuning (SFT) — supervised fine-tuning on thousands to millions of prompt–answer pairs written or checked by people, so the model answers instead of just continuing.
  3. Preference tuning (RLHF or DPO) — people or a reward model compare two answers; the model is trained to prefer the better one. This shapes helpfulness, tone and refusals.
  4. Reasoning RL (on many 2026 models) — reinforcement learning on problems with checkable answers (maths, code, tests), which teaches the model to "think" in hidden tokens before it answers.

The 2026 landscape in one paragraph

Frontier closed models (the GPT, Claude and Gemini families) are reached through APIs. Strong open-weight families — Llama, Qwen, DeepSeek, Mistral, Gemma and OpenAI's gpt-oss — can be downloaded and self-hosted. Most frontier models are now reasoning models with built-in thinking you control with an effort setting, many use mixture-of-experts layers, and context windows of about 1 million tokens exist on several of them.

What follows from the design

  • Probabilistic — the same prompt can give different answers.
  • Frozen knowledge — nothing after the training cutoff, unless you put it in the prompt.
  • No built-in fact check — a fluent sentence and a true sentence look the same to the model. This is the root of hallucination.

A real-life example

A bank puts an LLM behind its customer-support chat. A customer asks, "What is the interest rate on a 1-year fixed deposit?" The model replies, "The current rate is 7.1% per annum." It sounds right, but the model has no access to today's rate table — it produced a plausible number from patterns in its training data. The real rate was 6.8%.

The engineering team does not retrain the model. They fetch the live rate table from the bank's database and place it in the prompt, with the instruction "Answer only from the rates below; if the product is not listed, say you don't know." The model now copies 6.8% from context. The lesson: the model is a language engine, and facts that change must come from outside it.

Follow-up questions to expect

  • "What is the difference between a base model and a chat model?" — A base model only continues text; a chat model has been instruction-tuned and preference-tuned to follow instructions, hold a conversation and refuse unsafe requests.
  • "Why do LLMs hallucinate?" — Training rewards likely text, not true text. When the model lacks the fact, the most likely continuation is still a confident-sounding answer.
  • "What is a reasoning model?" — A model trained with reinforcement learning to generate hidden thinking tokens before the final answer. It is slower and costlier per query but much better at maths, code and multi-step problems.
  • "Is an LLM just autocomplete?" — Mechanically yes, it predicts the next token; but predicting well at scale requires learning a lot of structure, so it behaves far more capably than phone autocomplete.