LLMOps & Deployment

Course Content

LLMOps & Deployment

6 sections · 40 lessons

How does knowledge distillation create smaller, faster models?


What you need to know

Two kinds of distillation

  • Logit (soft-label) distillation — the student learns to match the teacher's probability over every next token. The relative probabilities ("refund" 0.6, "return" 0.3) teach which near-misses are acceptable. It needs access to the teacher's logits, so it is used when you own both models.
  • Sequence-level distillation — the student is fine-tuned on text the teacher wrote. This works with any teacher, including API models, and is what most application teams do.

The process

  1. Collect real inputs — thousands of production inputs covering all intents, languages and edge cases.
  2. Generate teacher outputs — with the best prompt you have, possibly with reasoning.
  3. Filter — keep outputs that pass evals, a verifier or human review; students copy mistakes faithfully.
  4. Fine-tune the student — usually LoRA on an open 7–8B model.
  5. Evaluate on a held-out set — including out-of-distribution cases.
  6. Serve cheaply — quantize and self-host, with the teacher as a fallback for hard cases.

What decides success

  • Coverage beats volume. 20,000 well-spread examples beat 200,000 from one narrow slice.
  • Narrow task. Distillation shines on one well-defined job (classify and summarise tickets), not on general chat.
  • Licences. Some API providers' terms restrict using outputs to train competing models; open-weight teachers have their own licence conditions. Check before you start.
  • A routing fallback. Low-confidence inputs still go to the teacher.

A real-life example

A fintech support bot hands complex chats to human agents with an auto-generated summary and intent label. This runs on a frontier API: 1.2 million handoffs and summaries a month, about 2,000 input and 250 output tokens each, at illustrative prices of $3/$15 per million — about $11,700 a month, with a p95 latency of 2.5 s.

The team distils it:

  • 30,000 real conversations, sampled across intents and languages, are sent to the teacher. Generating the data costs about $315.
  • Human reviewers check a 2,000-case sample; outputs failing the rubric are fixed or removed (about 6%).
  • An open 8B model is fine-tuned with LoRA on one GPU in a few hours.
  • On a 1,000-case held-out set, the student scores 94% of the teacher's rubric score, and slightly better on the label format.

Served in FP8 on one self-hosted GPU (roughly $1,800 a month, illustrative), p95 latency falls to 0.6 s. Conversations where the student's intent confidence is low (about 7%) still go to the teacher. The monthly bill drops by about 75%, and the same pipeline re-distils the student every quarter from fresh traffic.

Follow-up questions to expect

  • "Distillation or just prompt a small model?" — Try prompting first; distil when the small model's prompted quality is not enough and volume is high enough to pay back the work.
  • "How is this different from fine-tuning on human labels?" — The labels come from a stronger model instead of people, which is cheaper and faster but copies the teacher's errors, so filtering matters.
  • "What breaks first in production?" — New kinds of input the training data did not cover; monitor drift and keep the teacher fallback.