LLMs Deep Dive

Course Content

LLMs Deep Dive

10 sections · 40 lessons

What is model distillation, and how is it applied to LLMs?


What the student learns after 'Your refund has been'1.00.00.00.800.150.050.590.260.15processedinitiatedcreditedHard labelTeacher, T = 1Teacher, T = 2
The softened teacher tells the student that 'initiated' was nearly right — information the hard label throws away.

What you need to know

Hard labels versus soft labels

A bank bot's next word after "Your refund has been" — the training text says "processed".

Label typeprocessedinitiatedcredited
Hard label (the data)1.00.00.0
Teacher, T = 10.800.150.05
Teacher, T = 2 (softened)0.590.260.15

The hard label says only "processed is right". The teacher also says "initiated is a reasonable alternative, credited less so". That extra information — sometimes called dark knowledge — lets the student learn more from each example.

The classic loss

Text
loss = α · CE(student, hard label)     + (1 - α) · T² · KL(teacher_T || student_T)

teacher_T and student_T are softmax with temperature T; the T² keeps the gradient size comparable when T changes.

Three forms for LLMs

  1. Response (sequence-level) distillation — prompt the teacher, collect its answers, fine-tune the student on them. Only needs text, so it works with API teachers. Most small instruction-tuned models were built partly this way.
  2. Logit distillation — match the teacher's token distributions at each position. Needs the teacher's probabilities, so it works with open-weight teachers or APIs that return log-probabilities.
  3. Reasoning distillation — train on the teacher's worked solutions. DeepSeek, for example, released smaller Qwen- and Llama-based models distilled from its R1 reasoning model's outputs.

A newer variant, on-policy distillation, lets the student generate and has the teacher score or correct those outputs, so the student learns from its own mistakes.

Results and limits

The DistilBERT paper reported a model 40% smaller and 60% faster that kept about 97% of BERT's language-understanding performance. For LLMs, the typical pattern is a small student that matches the teacher on the narrow task it was distilled for, but not in general. Limits: the student copies teacher errors and biases, it cannot exceed the teacher on the distilled task by imitation alone, and many API providers' terms forbid using outputs to train competing models — check before you build.

A real-life example

An e-commerce search assistant uses a large frontier model to turn free-text queries into structured filters: category, brand, price range, size. It works well, but it costs about Rs 0.40 per query and takes 1.5 seconds; at 5 million queries a day, that is Rs 20 lakh a day (illustrative numbers).

The team collects 500,000 real queries, has the large model produce the structured filters, spot-checks 2,000 by hand, and fine-tunes a 1B open-weight student on the pairs. On a held-out set of 3,000 queries, the student matches the teacher on the common categories and is weaker on rare ones like "gifts for a Bengali wedding". They route queries the student is unsure about (low probability on its top output) to the large model — about 8% of traffic — and serve the rest from the student at a small fraction of the cost and under 100 ms.

Follow-up questions to expect

  • "Why use temperature in distillation?" — A sharp teacher distribution is almost one-hot; raising T spreads probability onto the alternatives so the student can learn from them.
  • "Distillation versus quantization?" — Distillation trains a new, smaller model; quantization stores the same model's weights in fewer bits. They are often combined.
  • "Can the student beat the teacher?" — On the distilled task it usually does not exceed the teacher, though filtering the teacher's outputs for correctness can give a cleaner dataset than the teacher's average behaviour.