Course Content
Fine-Tuning LLMs
6 sections · 52 lessons
Explain prefix tuning vs prompt tuning. How do they compare to LoRA?
What you need to know
Prompt tuning
Prepend N learned vectors to the embedded input. Only these are trained. On Llama-3.1-8B (embedding size 4,096), 20 virtual tokens are 20 × 4,096 = 81,920 parameters — about 160 KB in bf16.
Its original paper (Lester et al., 2021) found it becomes competitive with full fine-tuning only for very large models, and is clearly weaker on smaller ones.
Prefix tuning
Learn key and value vectors that are added at every attention layer, so every layer can attend to them. On Llama-3.1-8B, which has 32 layers and a key/value size of 1,024 (because of grouped-query attention):
20 tokens x 32 layers x 2 (key and value) x 1,024 = 1,310,720 parametersIt was introduced by Li and Liang (2021) and is usually trained through a small MLP for stability (prefix_projection=True in PEFT). P-tuning v2 is a similar "deep prompt" method.
Comparison
| Prompt tuning | Prefix tuning | LoRA (r=16, all-linear) | |
|---|---|---|---|
| Where it acts | Input embeddings | Keys and values at every layer | Weight matrices |
| Parameters on an 8B model | 82K | 1.3M | 41.9M |
| Quality on hard tasks | Weakest | Middle | Best |
| Merge into the model | No | No | Yes |
| Uses context / KV cache | Yes, N positions | Yes, N positions per layer | No |
| Mix tasks in one batch | Easy | Easy | Needs multi-LoRA serving |
Code (peft 0.21)
1from peft import PromptTuningConfig, get_peft_model23cfg = PromptTuningConfig(4 task_type="CAUSAL_LM", num_virtual_tokens=20,5 prompt_tuning_init="TEXT",6 prompt_tuning_init_text="Write a warm product description for this brand:",7 tokenizer_name_or_path="Qwen/Qwen2.5-7B-Instruct")8model = get_peft_model(base, cfg)Starting the soft prompt from the embeddings of a real sentence ("TEXT" init) trains faster and more stably than random vectors.
A real-life example
The description-writing company wants a custom voice for 5,000 small sellers, not just 40 big brands. Prompt tuning is tempting: 5,000 soft prompts at 160 KB each is only 800 MB, and requests for different sellers can share one batch.
On their 8B model, though, soft prompts capture little of each seller's voice. Brand reviewers rate them barely better than a written style description in the prompt. The team keeps LoRA adapters for the 40 big brands, and for the long tail uses a plain prompt that includes three of the seller's own past descriptions. Prompt tuning lost on quality at this model size.
Follow-up questions to expect
- "Why is prompt tuning weak on smaller models?" — Only a few input vectors steer the whole network; a smaller model has less capacity to be steered that way. Prefix tuning and LoRA act inside every layer.
- "Is prompt tuning the same as prompt engineering?" — No. Prompt engineering writes words by hand. Prompt tuning learns vectors with gradient descent, and they are not words.
- "When would you pick prompt or prefix tuning today?" — Thousands of tasks on one very large model where storage and mixed-task batching matter more than peak quality.