Course Content
Applied AI Engineering: From Prompt to Production
9 sections · 29 lessons
LoRA and parameter-efficient fine-tuning in practice
Full fine-tuning updates every weight in a model. With the usual Adam optimiser and mixed precision, you need roughly 16 bytes of GPU memory per parameter: the weights, their gradients and the optimiser's running statistics. For a 7-billion-parameter model that is about 112 GB before you store a single training example, which means several large GPUs. Even for a 0.5-billion-parameter model, it is about 8 GB, and every experiment produces another full copy of the model to store.
Parameter-efficient fine-tuning (PEFT) avoids this by freezing the original weights and training a small number of new ones. The most widely used method is LoRA, short for low-rank adaptation. With LoRA, PolicyPal's router trains about 0.2% of the model's parameters, fits comfortably on one small GPU, and saves an adapter of a few megabytes instead of a full model.
This lesson explains the idea, does the arithmetic on a real model, and trains the router.
The idea: a small correction to a big matrix
Most of a transformer's weights sit in large matrices. A layer's query projection in the model PolicyPal uses is a 896 × 896 matrix, about 800,000 numbers. Full fine-tuning learns a change to every one of them.
LoRA's insight is that the change needed to adapt a model to a narrow task is usually simple, and can be approximated by the product of two thin matrices. Instead of learning a full 896 × 896 update, LoRA learns a matrix A of size r × 896 and a matrix B of size 896 × r, where r, the rank, is small, often 8 or 16. The layer then computes with W + (α / r) × B × A, where W is frozen and α is a scaling setting.
With r = 16, the update has 16 × (896 + 896) = 28,672 trainable numbers instead of 802,816, about 3.6% of that matrix. B starts at zero, so at the first step the model behaves exactly like the original. Training only moves it as far as the data demands.
Counting the parameters for PolicyPal's router
PolicyPal's router starts from Qwen/Qwen2.5-0.5B, an open-weight model with about 494 million parameters. Its relevant shape: a hidden size of 896, 24 layers, and attention that uses 14 query heads but only 2 key-value heads of size 64. That last detail means the query projection maps 896 to 896, while the value projection maps 896 to only 2 × 64 = 128.
Apply LoRA with r = 16 to the query and value projections in every layer.
- Query: 16 × (896 + 896) = 28,672 per layer
- Value: 16 × (896 + 128) = 16,384 per layer
- Per layer: 45,056. Across 24 layers: 1,081,344
- The new classification head, trained fully: 896 × 6 labels = 5,376
- Total trainable: about 1.09 million, or about 0.22% of the model
Saved in 32-bit floats, that is about 4.4 MB. The frozen base model is loaded once and shared. If Harbourline later trains a second adapter for another task, it adds another few megabytes, not another full model.
Why a classification head and not generated text
You could fine-tune the model to write a label, such as "leave_balance", as text. For a closed label set there is a better way. AutoModelForSequenceClassification puts a small linear layer, the head, on top of the model's final hidden state and outputs one score per label in a single forward pass. There is no text to parse, no chance of an invented label, and you get a probability for every route. Those probabilities become important in the next lesson, where PolicyPal uses a threshold on the sensitive route.
The training script
The labelled data is in JSON Lines files with a text field and a label field. The next lesson shows how those files are built.
1# policypal/router/train.py2import numpy as np3from datasets import load_dataset4from peft import LoraConfig, TaskType, get_peft_model5from transformers import (AutoModelForSequenceClassification, AutoTokenizer,6 DataCollatorWithPadding, Trainer, TrainingArguments)78LABELS = ["policy_question", "leave_balance", "it_ticket",9 "multi_step", "sensitive_hr", "out_of_scope"]10BASE = "Qwen/Qwen2.5-0.5B"1112tok = AutoTokenizer.from_pretrained(BASE)13model = AutoModelForSequenceClassification.from_pretrained(14 BASE, num_labels=len(LABELS), id2label=dict(enumerate(LABELS)),15 label2id={name: i for i, name in enumerate(LABELS)})16model.config.pad_token_id = tok.pad_token_id # needed to find each row's last token1718model = get_peft_model(model, LoraConfig(task_type=TaskType.SEQ_CLS, r=16, lora_alpha=32,19 lora_dropout=0.05, target_modules=["q_proj", "v_proj"]))20model.print_trainable_parameters() # about 1.09M trainable, about 0.22%2122data = load_dataset("json", data_files={"train": "router/train.jsonl",23 "validation": "router/val.jsonl"})2425def encode(batch):26 enc = tok(batch["text"], truncation=True, max_length=128)27 enc["labels"] = [LABELS.index(name) for name in batch["label"]]28 return enc2930data = data.map(encode, batched=True, remove_columns=["text", "label"])3132def accuracy(pred):33 return {"accuracy": float((np.argmax(pred.predictions, axis=-1) == pred.label_ids).mean())}3435args = TrainingArguments(36 output_dir="router/out", num_train_epochs=3, learning_rate=2e-4,37 per_device_train_batch_size=16, per_device_eval_batch_size=64,38 eval_strategy="epoch", save_strategy="epoch", load_best_model_at_end=True,39 metric_for_best_model="accuracy", bf16=True, logging_steps=20, report_to="none")40trainer = Trainer(model=model, args=args, train_dataset=data["train"],41 eval_dataset=data["validation"], data_collator=DataCollatorWithPadding(tok),42 compute_metrics=accuracy)43trainer.train()44model.save_pretrained("router/adapter") # LoRA weights and the head: a few MB45tok.save_pretrained("router/adapter")A few lines deserve attention. Setting pad_token_id matters for this model family: the classification head reads the hidden state of each row's last real token, and it finds that token by looking for padding. With TaskType.SEQ_CLS, peft automatically keeps the new classification head trainable and saves it with the adapter. The learning rate of 2e-4 is about ten times higher than you would use for full fine-tuning, which is normal for LoRA because far fewer weights are moving. bf16=True needs a GPU with bfloat16 support, such as an L4, A10 or A100; on an older T4, use fp16=True instead.
On 2,400 training examples, three epochs took about 8 minutes on one L4 GPU. Validation accuracy was 94.3% after the first epoch and 95.9% after the third.
Choosing rank, targets and epochs
The team tried a few settings on the validation set. None of them changed much, which is typical for a narrow task.
| Setting | Trainable parameters | Validation accuracy |
|---|---|---|
| r = 8, query and value | 0.55M | 95.1% |
| r = 16, query and value | 1.09M | 95.9% |
| r = 32, query and value | 2.17M | 95.8% |
| r = 16, all linear layers | 8.8M | 96.1% |
Going to all linear layers gained 0.2 points for eight times the parameters, and the gain was inside the noise of a 300-example validation set. The team kept r = 16 on query and value. Watch the validation curve for overfitting: with small data, a fourth or fifth epoch often lowers training loss while validation accuracy stops rising or falls.
For larger models, QLoRA loads the frozen base in 4-bit precision with the bitsandbytes library and trains LoRA adapters on top. That brings a 7-billion-parameter fine-tune down to a single 24 GB GPU. PolicyPal's router does not need it; the 0.5B model fits easily as it is.
Serving the adapter
For inference, load the base model with the adapter and merge them into one set of weights, so there is no extra cost per call.
1# policypal/router/serve.py2import torch3from peft import PeftModel4from transformers import AutoModelForSequenceClassification, AutoTokenizer56LABELS = ["policy_question", "leave_balance", "it_ticket",7 "multi_step", "sensitive_hr", "out_of_scope"]89tok = AutoTokenizer.from_pretrained("router/adapter")10base = AutoModelForSequenceClassification.from_pretrained("Qwen/Qwen2.5-0.5B",11 num_labels=len(LABELS))12base.config.pad_token_id = tok.pad_token_id13model = PeftModel.from_pretrained(base, "router/adapter").merge_and_unload().eval()1415@torch.no_grad()16def route_probs(text: str) -> dict[str, float]:17 enc = tok(text, return_tensors="pt", truncation=True, max_length=128)18 probs = model(**enc).logits[0].softmax(-1).tolist()19 return dict(zip(LABELS, probs))The base model is created with the same six labels before the adapter is attached, so the saved classification head fits it exactly. merge_and_unload adds the LoRA matrices into the frozen weights and returns an ordinary model, so inference costs the same as the base model. On a small GPU, one message takes about 15 milliseconds; on a 4-core CPU, about 60. That is the number that replaces the prompted router's 850.
Check your understanding
0 of 3 answered
1.In Qwen2.5-0.5B, why does a rank-16 LoRA on the value projection have fewer parameters (16,384) than on the query projection (28,672)?
2.Why does PolicyPal train a classification head instead of fine-tuning the model to write the label as text?
3.The run with LoRA on all linear layers scored 96.1% on validation, against 95.9% for query and value only. Why did the team keep the smaller configuration?