Course Content
Fine-Tuning LLMs
6 sections · 52 lessons
How does knowledge distillation aid fine-tuning and what legal or usage limits should be considered?
What you need to know
Three kinds of distillation
| Kind | What the student learns from | Needs |
|---|---|---|
| Response (sequence-level) | Teacher's written answers, trained with plain SFT | Only text — works with API teachers |
| Logit (token-level) | Teacher's token distributions, with a KL loss and a temperature | Teacher weights or logits, shared tokenizer |
| On-policy | Teacher's distributions on answers the student wrote | Teacher running during training |
On-policy distillation fixes a weakness of the other two: in them, the student only sees the teacher's good answers, never its own mistakes. When the student writes the answers and the teacher scores every token, the student learns how to recover from its own errors. TRL 1.13 ships a DistillationTrainer for token-level distillation that can generate completions from the student during training.
Why it helps fine-tuning
A frontier model may be too costly, too slow or not allowed to see your data in production. Distillation moves its skill on your task into a small model you can host. The student will not match the teacher in general — only where you distilled.
Legal and usage limits
- API terms of service. Several major providers' terms forbid using their outputs to develop models that compete with them. Read your agreement before generating data, not after.
- Open-weight licences. They differ. For example, the Llama 3.1 community licence lets you use outputs to improve other models, but requires "Llama" at the start of the name of any model you distribute that was built that way. Apache-2.0 models have far fewer conditions.
- Data protection. Sending customer data to a teacher API is a transfer of personal data; in India that falls under the DPDP Act 2023, so check consent and contracts.
- Provenance. Record which teacher, which version and which terms applied for every dataset. "We distilled it" is something you may have to prove was allowed.
A real-life example
A hospital network wants a fast radiology summariser. Reports cannot leave its data centre, so the teacher is a large open-weight model it hosts itself, and the student is an 8B model.
The teacher writes summaries for 20,000 anonymised reports. Radiologists review 5% and reject about 6%, which leads the team to add a rule to the teacher prompt about always stating nodule size in millimetres. The student, trained with response distillation, runs about 8 times faster and scores close to the teacher on the radiologists' 400-report test set, while falling well behind it on general questions — as expected. The legal team checks the teacher's licence for naming and attribution terms before the model is shared with partner hospitals.
Follow-up questions to expect
- "Response or logit distillation — which do you use?" — Response distillation when the teacher is only available through an API or the tokenizers differ; logit or on-policy distillation when you have the teacher's weights and the same tokenizer, because it carries more information per example.
- "Does the student inherit the teacher's biases?" — Yes, and often amplified, because it has less capacity to learn nuance. Evaluate bias and safety on the student separately.
- "What temperature do you use?" — In logit distillation, a temperature above 1 softens both distributions so the student learns from the teacher's second and third choices; 1–2 is common.