Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

Explain prefix tuning vs prompt tuning. How do they compare to LoRA?


Trainable parameters on Llama-3.1-8BInput embeddings81,920NoK and V,every layer1,310,720NoWeight matrices41,943,040YesActs onParametersMergeablePrompt tuning, 20 tokensPrefix tuning, 20 tokensLoRA r=16
Prompt and prefix tuning are hundreds of times smaller, but they steer from outside the weights — which is why they lose on quality below very large model sizes.

What you need to know

Prompt tuning

Prepend N learned vectors to the embedded input. Only these are trained. On Llama-3.1-8B (embedding size 4,096), 20 virtual tokens are 20 × 4,096 = 81,920 parameters — about 160 KB in bf16.

Its original paper (Lester et al., 2021) found it becomes competitive with full fine-tuning only for very large models, and is clearly weaker on smaller ones.

Prefix tuning

Learn key and value vectors that are added at every attention layer, so every layer can attend to them. On Llama-3.1-8B, which has 32 layers and a key/value size of 1,024 (because of grouped-query attention):

Text
20 tokens x 32 layers x 2 (key and value) x 1,024 = 1,310,720 parameters

It was introduced by Li and Liang (2021) and is usually trained through a small MLP for stability (prefix_projection=True in PEFT). P-tuning v2 is a similar "deep prompt" method.

Comparison

Prompt tuningPrefix tuningLoRA (r=16, all-linear)
Where it actsInput embeddingsKeys and values at every layerWeight matrices
Parameters on an 8B model82K1.3M41.9M
Quality on hard tasksWeakestMiddleBest
Merge into the modelNoNoYes
Uses context / KV cacheYes, N positionsYes, N positions per layerNo
Mix tasks in one batchEasyEasyNeeds multi-LoRA serving

Code (peft 0.21)

Python
from peft import PromptTuningConfig, get_peft_modelcfg = PromptTuningConfig(    task_type="CAUSAL_LM", num_virtual_tokens=20,    prompt_tuning_init="TEXT",    prompt_tuning_init_text="Write a warm product description for this brand:",    tokenizer_name_or_path="Qwen/Qwen2.5-7B-Instruct")model = get_peft_model(base, cfg)

Starting the soft prompt from the embeddings of a real sentence ("TEXT" init) trains faster and more stably than random vectors.

A real-life example

The description-writing company wants a custom voice for 5,000 small sellers, not just 40 big brands. Prompt tuning is tempting: 5,000 soft prompts at 160 KB each is only 800 MB, and requests for different sellers can share one batch.

On their 8B model, though, soft prompts capture little of each seller's voice. Brand reviewers rate them barely better than a written style description in the prompt. The team keeps LoRA adapters for the 40 big brands, and for the long tail uses a plain prompt that includes three of the seller's own past descriptions. Prompt tuning lost on quality at this model size.

Follow-up questions to expect

  • "Why is prompt tuning weak on smaller models?" — Only a few input vectors steer the whole network; a smaller model has less capacity to be steered that way. Prefix tuning and LoRA act inside every layer.
  • "Is prompt tuning the same as prompt engineering?" — No. Prompt engineering writes words by hand. Prompt tuning learns vectors with gradient descent, and they are not words.
  • "When would you pick prompt or prefix tuning today?" — Thousands of tasks on one very large model where storage and mixed-task batching matter more than peak quality.