Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

For a real product, how do you choose among prompt tuning, LoRA, distillation, and RLHF?


What you need to know

The four techniques

  • Prompt tuning — learns a few "soft prompt" vectors that are added to the input, while the model stays frozen. It is not the same as prompt engineering (writing better text prompts). It is tiny and cheap, but on today's models it usually underperforms LoRA (see Section 2).
  • LoRA — trains small low-rank matrices inside the model. The standard way to teach a task, format or tone from examples.
  • Distillation — trains a smaller student model on a larger teacher's outputs (or probabilities), to get similar quality at lower cost. Check the teacher's terms of use (Section 2).
  • RLHF / preference tuning — trains on "better versus worse" signals. DPO does this without a reward model; PPO and GRPO use a reward.

Map the symptom to the technique

SymptomReach for
Format or style inconsistent, little dataBetter prompting; then prompt tuning or a small LoRA
Task accuracy too low, thousands of examplesLoRA
Quality fine, but too slow or costlyDistillation into a smaller model
Correct but verbose, unhelpful or unsafeDPO, then RLHF if needed
Facts wrong or change oftenNone of these — RAG

The order that usually works

  1. Prompting and RAG — establish the ceiling and, most importantly, build the eval set.
  2. LoRA on curated data — the cheapest way to change weights.
  3. DPO — once production feedback arrives as preferences (edits, thumbs, A/B picks).
  4. Distillation — last, as a cost step, once you know exactly what behaviour to preserve.

What to measure at each step

Offline: task accuracy, format validity, and a general-capability regression suite. Online: task success rate, escalation rate, p95 latency and cost per request.

A real-life example

An e-commerce company needs a brand-voice product-description writer for 2 million products.

  • Week 1 — prompting. A large hosted model with a style guide and 5 examples. The brand team approves 78% of drafts. They save 500 approved descriptions as the eval set.
  • Week 3 — LoRA. A LoRA on an 8B open model, trained on 3,000 approved descriptions. Approval reaches 84%, and the model runs on their own GPUs.
  • Week 7 — DPO. The team's A/B picks between two drafts become 4,000 preference pairs. DPO cuts overused phrases, and approval reaches 90%.
  • Week 10 — distillation. For the catalogue backfill, the 8B model generates 60,000 descriptions and a 3B student is trained on them. The student scores 87% at about a third of the serving cost, so it handles bulk backfill while the 8B handles new launches.

(Numbers are from this made-up scenario.)

Follow-up questions to expect

  • "When would you use full RLHF instead of DPO?" — When you have a reliable reward (a trusted reward model or verifiable checks) and need the model to improve beyond the answers in your pairs. Otherwise DPO is simpler and cheaper.
  • "Can you combine them?" — Yes, and most real systems do: RAG for facts, LoRA for task and format, DPO for preferences, distillation for cost.
  • "Why not fine-tune from day one?" — Without an eval set you cannot tell whether fine-tuning helped. Prompting builds that set cheaply.