Course Content
Fine-Tuning LLMs with LoRA, QLoRA and PEFT
4 sections · 10 lessons
Training with QLoRA — An End-to-End PEFT Pipeline
Here is a training script that will run to completion, report a beautifully decreasing loss, and produce a model that has learned nothing useful:
1texts = [ex["instruction"] + ex["output"] for ex in dataset]2tokens = tokenizer(texts, truncation=True, max_length=512)3trainer = Trainer(model=model, train_dataset=tokens, args=TrainingArguments(...))4trainer.train()Four defects, none of which raise an error. There is no separator between instruction and output, so the model never learns where its answer should begin. There is no end-of-sequence token, so at inference it generates until it hits the token limit. Loss is computed over the instruction tokens as well as the answer, so most of the gradient signal is spent teaching the model to predict questions. And max_length=512 silently truncates 30% of the examples mid-answer, teaching the model to stop in the middle of sentences.
Every one of those produces a smooth loss curve. Fine-tuning goes wrong quietly, and the defects live in data preparation far more often than in the training loop. This is the full pipeline with those traps closed.
The environment
1pip install "torch>=2.6" \2 "transformers>=5.0" \3 "peft>=0.21" \4 "bitsandbytes>=0.45" \5 "accelerate>=1.0" \6 "datasets>=3.0" \7 "trl>=1.0"The code in this course was checked against transformers 5.17, peft 0.21, trl 1.13 and bitsandbytes 0.50. These libraries rename arguments between major versions — transformers 5 uses dtype= instead of torch_dtype=, and replaced warmup_ratio with a fractional warmup_steps — so if a tutorial's code fails on an unexpected keyword, check its versions against yours first.
Check the GPU before anything else, because the answer changes your configuration:
1import torch23print(torch.cuda.get_device_name(0))4print(f"{torch.cuda.get_device_properties(0).total_memory / 1e9:.1f} GB")5print("bf16 supported:", torch.cuda.is_bf16_supported())If is_bf16_supported() is False — a T4 or V100 — you must use FP16 everywhere and accept that you may need loss scaling. On an A10, A100, L4 or any RTX 30/40-series card, use BF16 and skip that whole class of problem.
Preparing the data
Look at it first
1from datasets import load_dataset23ds = load_dataset("json", data_files="data/support.jsonl", split="train")4print(ds)5print(ds[0])6# {'instruction': 'Classify this support ticket and draft a reply.',7# 'input': 'my order 88213 said delivered tuesday but nothings here',8# 'output': 'Category: delivery_issue\n\nHi, ...'}Print three examples in full. Read them. This is not ceremony — it is where you catch trailing agent signatures, HTML entities that survived the export, and empty outputs.
Formatting into a single training string
A causal language model trains on one flat sequence of tokens. Your prompt and completion must be joined into that sequence with explicit, consistent markers, and the same template must be used at inference or the model will not recognise its own format.
1PROMPT_TEMPLATE = (2 "### Instruction:\n{instruction}\n\n"3 "### Input:\n{input}\n\n"4 "### Response:\n"5)67def format_example(ex):8 prompt = PROMPT_TEMPLATE.format(instruction=ex["instruction"],9 input=ex.get("input", ""))10 # The EOS token is what teaches the model to stop generating.11 return {"prompt": prompt, "full": prompt + ex["output"] + tokenizer.eos_token}1213ds = ds.map(format_example)14print(repr(ds[0]["full"][-80:]))15# '...refunded within 3 working days.</s>' <- confirm EOS is really thereForgetting the EOS token is the single most common cause of a fine-tuned model that rambles forever. It is one string concatenation and it is invisible in the loss curve.
Tokenising, and masking the prompt
By default, a causal LM computes loss on every token, including the instruction. That is not what you want. You want the model to learn the answer given the question, not to learn to generate questions. Setting label positions to -100 tells PyTorch's cross-entropy to ignore them.
1MAX_LEN = 102423def tokenize(ex):4 full = tokenizer(ex["full"], truncation=True, max_length=MAX_LEN)5 prompt_len = len(tokenizer(ex["prompt"], truncation=True,6 max_length=MAX_LEN)["input_ids"])78 labels = list(full["input_ids"])9 labels[:prompt_len] = [-100] * prompt_len # ignore the prompt10 full["labels"] = labels11 return full1213tokenized = ds.map(tokenize, remove_columns=ds.column_names)Masking typically changes the reported loss noticeably — it often rises, because the easy, highly-predictable template tokens no longer contribute. That is a good sign, not a regression: the number is now measuring the thing you care about.
Choosing max_length from the data, not from habit
1lengths = sorted(len(tokenizer(x["full"])["input_ids"]) for x in ds)2n = len(lengths)3print("p50", lengths[n // 2],4 "p90", lengths[int(n * 0.90)],5 "p99", lengths[int(n * 0.99)],6 "max", lengths[-1])7# p50 384 p90 812 p99 1436 max 3902With that distribution, MAX_LEN=1024 truncates about 6% of examples and MAX_LEN=512 truncates about 30%. Since memory scales with sequence length, the honest trade is: set the limit near p95, and drop the handful of examples that exceed it rather than truncating them mid-answer. A truncated answer is a training example that teaches the model to stop early.
Splitting
split = tokenized.train_test_split(test_size=0.1, seed=42)train_ds, eval_ds = split["train"], split["test"]print(len(train_ds), len(eval_ds)) # 2700 300Fix the seed. An unseeded split means every rerun evaluates on different data and your run-to-run comparisons are meaningless.
Loading the base model in 4 bits
1import torch2from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig34MODEL = "meta-llama/Llama-2-7b-hf"56bnb = BitsAndBytesConfig(7 load_in_4bit=True,8 bnb_4bit_quant_type="nf4",9 bnb_4bit_use_double_quant=True,10 bnb_4bit_compute_dtype=torch.bfloat16,11)1213tokenizer = AutoTokenizer.from_pretrained(MODEL)14tokenizer.pad_token = tokenizer.eos_token15tokenizer.padding_side = "right" # "left" only for batched generation1617model = AutoModelForCausalLM.from_pretrained(18 MODEL, quantization_config=bnb, device_map="auto",19 dtype=torch.bfloat16,20)21model.config.use_cache = False # incompatible with gradient checkpointing2223print(f"{model.get_memory_footprint() / 1e9:.2f} GB") # ~3.9 GBTwo lines there are load-bearing and frequently omitted. pad_token = eos_token is needed because Llama ships without a pad token and the collator will crash without one. use_cache = False is needed because the key/value cache conflicts with gradient checkpointing; leaving it True produces a warning you will ignore and a memory profile you will not understand.
Attaching the adapter
1from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training23model = prepare_model_for_kbit_training(model, use_gradient_checkpointing=True)45lora = LoraConfig(6 r=16,7 lora_alpha=32, # scale = 32/16 = 2.08 lora_dropout=0.05,9 bias="none",10 task_type="CAUSAL_LM",11 target_modules=["q_proj", "k_proj", "v_proj", "o_proj",12 "gate_proj", "up_proj", "down_proj"],13)1415model = get_peft_model(model, lora)16model.print_trainable_parameters()17# trainable params: 39,976,960 || all params: 6,778,392,576 || trainable%: 0.5898Verify that number against arithmetic you do yourself. For each target matrix of shape d×k the adapter adds r(d+k) parameters. In a 32-layer Llama-2-7B: the four attention projections are 4096 × 4096 each, contributing 4×16×8192=524,288 per layer; the three feed-forward matrices pair 11008 with 4096, contributing 3×16×15,104=724,992 per layer. That is 1,249,280 per layer, times 32 layers, equals 39,976,960. It matches, so every module was found. (The "all params" figure counts each 4-bit weight as one parameter even though two are packed into each byte, so it matches the BF16 model plus the adapter.)
If your printed count is smaller than your arithmetic, a name in target_modules does not exist in this architecture. PEFT raises an error only when no name matches; one wrong name among several is skipped silently, and training proceeds happily with fewer adapters than you think. List the real names with:
names = {n.split(".")[-1] for n, m in model.named_modules() if "Linear" in type(m).__name__}print(sorted(names))Training configuration
The collator
1from transformers import DataCollatorForSeq2Seq23collator = DataCollatorForSeq2Seq(4 tokenizer, padding=True, label_pad_token_id=-100, return_tensors="pt"5)This one pads labels with -100, which matters: the common alternative, DataCollatorForLanguageModeling(mlm=False), overwrites your carefully masked labels by copying input_ids. If you masked the prompt and then used that collator, the mask is silently discarded.
Training arguments
1from transformers import TrainingArguments23args = TrainingArguments(4 output_dir="runs/support-qlora",5 num_train_epochs=3,6 per_device_train_batch_size=4,7 gradient_accumulation_steps=4, # effective batch = 4 x 4 = 168 gradient_checkpointing=True,9 learning_rate=2e-4,10 lr_scheduler_type="cosine",11 warmup_steps=0.03, # a float below 1 is a fraction of all steps12 weight_decay=0.01,13 max_grad_norm=0.3,14 optim="paged_adamw_8bit",15 bf16=True,16 logging_steps=10,17 eval_strategy="steps",18 eval_steps=50,19 save_strategy="steps",20 save_steps=50,21 save_total_limit=3,22 load_best_model_at_end=True,23 metric_for_best_model="eval_loss",24 report_to="none",25 seed=42,26)Work out what this schedule actually does. 2,700 training examples with an effective batch of 16 gives 2700/16≈169 optimiser steps per epoch, so three epochs is roughly 507 steps. Warmup at 3% is about 15 steps. At a measured 3.2 seconds per step on an A10, the run takes 507 × 3.2 ≈ 1,620 seconds, or 27 minutes. Compute this before you launch. If your estimate says 40 hours, you want to know now.
Two settings deserve explanation. optim="paged_adamw_8bit" stores the optimiser moments in 8 bits rather than 32 and pages them to CPU memory under pressure — for 40M trainable parameters it saves about 240 MB and, more importantly, absorbs spikes. max_grad_norm=0.3 is tighter than the usual 1.0 because QLoRA runs are more prone to a single bad batch producing a large gradient; clipping harder is cheap insurance.
The trainer
1from transformers import Trainer23trainer = Trainer(4 model=model,5 args=args,6 train_dataset=train_ds,7 eval_dataset=eval_ds,8 data_collator=collator,9)1011trainer.train()12trainer.save_model("adapters/support-qlora") # writes ~160 MB, not 13.5 GB13tokenizer.save_pretrained("adapters/support-qlora")What lands on disk is adapter_model.safetensors plus adapter_config.json. The base model is not copied. Save the tokeniser alongside it — if you later add special tokens and forget, the adapter will be applied to a differently-tokenised input and behave strangely.
The same run with TRL
Everything above is done by hand so you can see it. In practice, most teams hand the formatting, masking and EOS handling to TRL's SFTTrainer. Give it a dataset of prompt and completion columns and it tokenises both, appends the EOS token to the completion, and computes loss on the completion only:
1from trl import SFTConfig, SFTTrainer23pairs = ds.map(lambda ex: {"prompt": PROMPT_TEMPLATE.format(4 instruction=ex["instruction"], input=ex.get("input", "")),5 "completion": ex["output"]}, # no EOS: TRL adds it6 remove_columns=ds.column_names)7pairs = pairs.train_test_split(test_size=0.1, seed=42)89sft_args = SFTConfig(10 output_dir="runs/support-qlora-trl",11 max_length=1024, # longer examples are truncated, not dropped12 completion_only_loss=True, # the default for prompt/completion data13 num_train_epochs=3, per_device_train_batch_size=4,14 gradient_accumulation_steps=4, learning_rate=2e-4,15 lr_scheduler_type="cosine", warmup_steps=0.03,16 optim="paged_adamw_8bit", bf16=True, max_grad_norm=0.3,17 eval_strategy="steps", eval_steps=50, save_steps=50, report_to="none",18)19trainer = SFTTrainer(model=model_4bit, args=sft_args, peft_config=lora,20 train_dataset=pairs["train"], eval_dataset=pairs["test"],21 processing_class=tokenizer)22trainer.train()Here ds is the dataset loaded under "Look at it first" (any columns besides instruction, input and output are dropped by the map), and model_4bit is the model straight out of from_pretrained with the BitsAndBytesConfig. Do not call get_peft_model yourself; pass the LoraConfig and SFTTrainer attaches the adapter. For chat data stored as a messages list, use assistant_only_loss=True instead, which trains only on the assistant turns. Older tutorials use DataCollatorForCompletionOnlyLM for this. It has been removed from TRL 1.x, so use these two flags instead.
The manual version is still worth knowing. When a TRL run misbehaves, the checks are the same: confirm that the EOS token is really at the end of the completion, and that the prompt positions really are -100 in trainer.train_dataset[0]["labels"].
Reading the training run
Three signals tell you almost everything.
| What you see | What it means | What to do |
|---|---|---|
| Train loss falls, eval loss falls, gap stays small | Healthy | Let it run |
| Train loss falls, eval loss flattens then rises | Overfitting from that step onwards | Stop at the eval minimum; fewer epochs or lower rank next time |
| Both losses flat from the start | Nothing is learning | Check trainable params > 0; raise LR; check labels are not all -100 |
| Loss spikes to NaN | Numerical blow-up | Lower LR; BF16 instead of FP16; tighten grad clipping |
| Loss decreases then jumps at each epoch boundary | Data ordering artefact | Confirm shuffling is on |
| Grad norm at exactly max_grad_norm every step | Clipping constantly; LR too high | Reduce LR by 3x |
Typical healthy numbers for an instruction-tuning run on a 7B base: initial loss around 1.8-2.2, settling to 0.8-1.2 by the end. A loss below 0.3 on a small dataset almost always means memorisation — check the eval gap.
1import matplotlib.pyplot as plt23hist = trainer.state.log_history4tr = [(h["step"], h["loss"]) for h in hist if "loss" in h]5ev = [(h["step"], h["eval_loss"]) for h in hist if "eval_loss" in h]67plt.plot(*zip(*tr), label="train")8plt.plot(*zip(*ev), label="eval", marker="o")9plt.xlabel("step"); plt.ylabel("loss"); plt.legend()10plt.savefig("runs/support-qlora/loss.png", dpi=120)Reloading and merging
For further training or evaluation
1from peft import PeftModel23base = AutoModelForCausalLM.from_pretrained(4 MODEL, quantization_config=bnb, device_map="auto")5model = PeftModel.from_pretrained(base, "adapters/support-qlora")6model.eval()For deployment
Do not merge into the 4-bit base. The merged weights would be re-rounded to 16 levels and quality drops measurably. Reload the base in BF16, apply the adapter there, then merge:
1base_fp = AutoModelForCausalLM.from_pretrained(2 MODEL, dtype=torch.bfloat16, device_map="cpu") # ~13.5 GB of RAM3merged = PeftModel.from_pretrained(base_fp, "adapters/support-qlora")4merged = merged.merge_and_unload() # W <- W + (alpha/r) B A5merged.save_pretrained("models/support-7b-merged") # safetensors by default6tokenizer.save_pretrained("models/support-7b-merged")The output is a plain 13.5 GB Llama checkpoint that any inference server loads without PEFT installed, and with zero adapter overhead at runtime. Keep the 160 MB adapter too — it is the artefact you can version, diff and re-merge later.
Sanity-check generation before you celebrate
1prompt = PROMPT_TEMPLATE.format(2 instruction="Classify this support ticket and draft a reply.",3 input="charged twice for order 44120, need the second one back")45ids = tokenizer(prompt, return_tensors="pt").to(model.device)6out = model.generate(**ids, max_new_tokens=200, do_sample=False,7 eos_token_id=tokenizer.eos_token_id)8print(tokenizer.decode(out[0][ids["input_ids"].shape[1]:],9 skip_special_tokens=True))Use greedy decoding (do_sample=False) for this check so the output is deterministic and comparable across runs. Confirm three things: the format matches your training targets, the model stops on its own, and the content is plausible. If it runs to max_new_tokens every time, your EOS token never made it into the training data.
Troubleshooting
Out of memory
Work down this list in order — each step costs less quality than the one after it:
| Change | Memory saved | Cost |
|---|---|---|
gradient_checkpointing=True | Large — often 60-70% of activation memory | ~30% slower per step |
Halve per_device_train_batch_size, double gradient_accumulation_steps | Proportional to batch | Slightly slower; effective batch unchanged |
Reduce MAX_LEN from 1024 to 512 | Roughly halves activations | Truncates long examples — check the distribution first |
optim="paged_adamw_8bit" | ~240 MB at 40M trainable params, plus spike absorption | Negligible |
| Target fewer modules (q, v only) | Cuts trainable params from 40M to 8.4M | Less adaptation capacity |
| Lower rank | Linear in r | Less capacity |
The activation arithmetic explains why checkpointing dominates. One hidden-state tensor at batch 4, sequence 1024, hidden 4096 in BF16 is 4×1024×4096×2=33.5 MB. A transformer block saves roughly 16 such tensors for the backward pass, so 32 layers hold about 16×33.5×32≈17 GB. Gradient checkpointing stores only one tensor per block boundary — about 1.07 GB — and recomputes the rest during backward. That is the trade: 30% more compute for 16 GB.
The model trains but the output is poor
| Symptom | Likely cause | Fix |
|---|---|---|
| Generates until the token limit, never stops | No EOS token in training targets | Append tokenizer.eos_token when formatting |
| Repeats the prompt back | Loss not masked; model learned to predict instructions | Set prompt label positions to -100 |
| Ignores the instruction entirely | Inference prompt template differs from training | Use the identical template string |
| Right content, wrong format | Too few examples, or rank too low | More examples; raise r from 8 to 16 or 32 |
| Perfect on training examples, poor otherwise | Overfitting | Fewer epochs; more dropout; more data |
| Cuts off mid-sentence | Training examples truncated by max_length | Raise MAX_LEN or drop over-long examples |
| Worse than the base model at everything | Learning rate far too high | Drop from 2e-4 to 5e-5 and rerun |
Training is slower than expected
Check, in this order: is gradient_checkpointing on when you did not need it (30% cost)? Is bnb_4bit_compute_dtype still at the FP32 default (roughly 2× cost)? Is padding dominating — a batch of one 1,500-token example and three 200-token examples pads everything to 1,500, wasting 70% of the compute? Sorting examples by length into buckets, or using example packing, recovers most of that.
Almost every "the fine-tune did not work" report resolves to one of four causes: no EOS token, unmasked prompt loss, a mismatched inference template, or a learning rate off by an order of magnitude. Check those four before you touch anything else.
What this means when you run a real job
Before the full run, do a deliberate 20-step smoke test: max_steps=20, eval_steps=10, on 200 examples. It takes about a minute and proves that the data loads, the collator produces the right shapes, the adapter has non-zero trainable parameters, memory fits, and the loss moves. Every failure described above except overfitting shows up in that minute.
Then, once the real run finishes, generate on ten held-out prompts and read the outputs yourself before you look at any metric. The loss curve cannot tell you that the model is politely answering a different question than the one you asked. A human reading ten outputs can tell you in two minutes, and that check has caught more broken fine-tunes than every automated metric combined.