Course Content
LLMOps & Deployment
6 sections · 40 lessons
How does knowledge distillation create smaller, faster models?
What you need to know
Two kinds of distillation
- Logit (soft-label) distillation — the student learns to match the teacher's probability over every next token. The relative probabilities ("refund" 0.6, "return" 0.3) teach which near-misses are acceptable. It needs access to the teacher's logits, so it is used when you own both models.
- Sequence-level distillation — the student is fine-tuned on text the teacher wrote. This works with any teacher, including API models, and is what most application teams do.
The process
- Collect real inputs — thousands of production inputs covering all intents, languages and edge cases.
- Generate teacher outputs — with the best prompt you have, possibly with reasoning.
- Filter — keep outputs that pass evals, a verifier or human review; students copy mistakes faithfully.
- Fine-tune the student — usually LoRA on an open 7–8B model.
- Evaluate on a held-out set — including out-of-distribution cases.
- Serve cheaply — quantize and self-host, with the teacher as a fallback for hard cases.
What decides success
- Coverage beats volume. 20,000 well-spread examples beat 200,000 from one narrow slice.
- Narrow task. Distillation shines on one well-defined job (classify and summarise tickets), not on general chat.
- Licences. Some API providers' terms restrict using outputs to train competing models; open-weight teachers have their own licence conditions. Check before you start.
- A routing fallback. Low-confidence inputs still go to the teacher.
A real-life example
A fintech support bot hands complex chats to human agents with an auto-generated summary and intent label. This runs on a frontier API: 1.2 million handoffs and summaries a month, about 2,000 input and 250 output tokens each, at illustrative prices of $3/$15 per million — about $11,700 a month, with a p95 latency of 2.5 s.
The team distils it:
- 30,000 real conversations, sampled across intents and languages, are sent to the teacher. Generating the data costs about $315.
- Human reviewers check a 2,000-case sample; outputs failing the rubric are fixed or removed (about 6%).
- An open 8B model is fine-tuned with LoRA on one GPU in a few hours.
- On a 1,000-case held-out set, the student scores 94% of the teacher's rubric score, and slightly better on the label format.
Served in FP8 on one self-hosted GPU (roughly $1,800 a month, illustrative), p95 latency falls to 0.6 s. Conversations where the student's intent confidence is low (about 7%) still go to the teacher. The monthly bill drops by about 75%, and the same pipeline re-distils the student every quarter from fresh traffic.
Follow-up questions to expect
- "Distillation or just prompt a small model?" — Try prompting first; distil when the small model's prompted quality is not enough and volume is high enough to pay back the work.
- "How is this different from fine-tuning on human labels?" — The labels come from a stronger model instead of people, which is cheaper and faster but copies the teacher's errors, so filtering matters.
- "What breaks first in production?" — New kinds of input the training data did not cover; monitor drift and keep the teacher fallback.