Course Content
Fine-Tuning LLMs with LoRA, QLoRA and PEFT
4 sections · 10 lessons
Hyperparameters and Checkpoints — Tuning a Run
A run has been going for six hours. Loss started at 2.1, fell steadily to 0.91, and at step 1,840 the pod is evicted. There are no checkpoints, because save_strategy was left at its default of "steps" with save_steps=500 and someone changed it to "no" a week ago to save disk. Six hours of A100 time, gone, with nothing to show for it — not even the loss history, which lived in the process.
A second team avoids that trap but hits a different one. Their loss curve looks fine, and the finished model is worse than the base model at everything. The cause: learning_rate=2e-3, ten times too high. The loss decreased because the model learned to reproduce the training set's surface patterns while its general capabilities collapsed. The curve told them nothing.
Hyperparameters decide whether a run learns the right thing; checkpoints decide whether you keep it. Both are cheap to get right and expensive to get wrong.
What you are actually tuning
| Parameter | Typical range (QLoRA, 7B) | Impact | Tune it? |
|---|---|---|---|
| learning_rate | 5e-5 to 5e-4 | Very high | Always — first thing |
| lora r | 8 to 64 | High | Usually — sweep 3 values |
| lora_alpha | 2r as a default | Medium | Tie to r; sweep rarely |
| num_train_epochs | 1 to 5 | High | Use early stopping instead |
| effective batch size | 8 to 128 | Medium | Set by memory, then match LR |
| warmup_steps (as a fraction) | 0.03 to 0.10 | Low-medium | Rarely — 0.03 works |
| lora_dropout | 0.0 to 0.1 | Low-medium | Raise if overfitting |
| weight_decay | 0.0 to 0.1 | Low | 0.01 and move on |
| max_grad_norm | 0.3 to 1.0 | Low, until it is not | Lower it if you see spikes |
| target_modules | 2 or 7 modules | High | Yes — often beats raising r |
Spend your sweep budget on the top of that table. Learning rate and rank between them explain most of the variance between a good and a bad run; weight decay explains almost none.
Effective batch size
Three separate settings combine into the number that actually matters:
Gradient accumulation runs several forward/backward passes, sums the gradients, and only then takes one optimiser step. Mathematically this is close to a single large batch, but peak memory is set by the per-device batch alone. That is the lever: you can train with an effective batch of 64 on a card that can only hold 4 sequences at a time.
per_device=16, accum=1 -> effective 16, memory for 16 sequences (may OOM)per_device=4, accum=4 -> effective 16, memory for 4 sequences (fits)per_device=1, accum=16 -> effective 16, memory for 1 sequence (fits anywhere)The cost is throughput. Sixteen separate forward passes have more per-step overhead than one batched pass, so per_device=1, accum=16 might run 30-40% slower per optimiser step than per_device=4, accum=4. Use the largest per-device batch that fits, then make up the rest with accumulation.
What batch size does to learning
- Small effective batch (4-8): noisy gradients. The noise acts as a regulariser and can help on small datasets, but training is unstable and one bad example can move the weights a long way.
- Medium (16-32): the sweet spot for most instruction fine-tuning. Enough averaging to be stable, enough noise to generalise.
- Large (64-256): smooth, stable, fast in wall-clock terms if you have the memory — but each step sees the same data averaged more heavily, so you need proportionally fewer steps and often a higher learning rate to compensate.
The rough correspondence is the square-root rule: if you multiply batch size by 4, multiply learning rate by about 2. It is a heuristic, not a law, but it stops the common mistake of quadrupling the batch to fix an OOM and then wondering why the model underfits.
| GPU | Memory | Model / precision | Suggested per_device × accum | Max length |
|---|---|---|---|---|
| T4 | 16 GB | 7B, NF4, fp16 | 1 × 16 | 512 |
| RTX 4090 | 24 GB | 7B, NF4, bf16 | 4 × 4 | 1024 |
| A10G | 24 GB | 7B, NF4, bf16 | 4 × 4 | 1024 |
| A100 | 40 GB | 13B, NF4, bf16 | 4 × 4 | 2048 |
| A100 | 80 GB | 70B, NF4, bf16 | 2 × 8 | 2048 |
Learning rate
Why LoRA wants a bigger one than you expect
Full fine-tuning of a 7B model typically uses 1e-5 to 5e-5. LoRA typically uses 1e-4 to 3e-4 — five to ten times higher. The reason is that LoRA's matrix B starts at exactly zero, so the adapter contributes nothing at step 0 and has to travel a meaningful distance before it does anything at all. There is also no risk of wrecking the pretrained weights, because they are frozen. A learning rate that would destroy a full fine-tune is merely brisk for an adapter.
Schedules
A constant learning rate is wrong at both ends of training. At the start, the adapter is random and the first few gradients are large and unreliable; a full-size step here can knock the run into a bad region it never recovers from. At the end, you want small refinements, and a full-size step overshoots the minimum repeatedly.
Warmup plus cosine decay fixes both:
During warmup, for step t≤Tw:
After warmup, with T total steps:
η(t)=2ηmax(1+cos(π⋅T−Twt−Tw))Worked through, with ηmax=2×10−4, T=375 total steps and Tw=30 warmup steps:
Step Progress through decay cos term Learning rate 0 — — 0 15 half of warmup — 1.00e-4 30 warmup complete 1.000 2.00e-4 100 70/345 = 0.203 0.803 1.80e-4 200 170/345 = 0.493 0.023 1.02e-4 300 270/345 = 0.783 -0.776 2.24e-5 375 1.000 -1.000 0 Notice the shape: the rate stays near its maximum for the first third, then falls away sharply. That is deliberate — most of the learning happens early, and the long tail of small steps is what makes the final model stable.
Schedule Behaviour Use when constant Flat throughout Debugging only constant_with_warmup Ramp then flat Continuing a run; unknown total length linear Ramp then straight decline to 0 Simple, predictable, fine cosine Ramp then smooth decline to 0 Default for fine-tuning cosine_with_restarts Repeated decay cycles Long runs escaping plateaus Finding the rate
Do not guess and do not sweep blindly. Run a short LR range test: 100 steps, increasing the rate exponentially, and plot loss against rate.
Python1import math, torch23def lr_range_test(model, loader, opt, lo=1e-6, hi=1e-2, steps=100):4 mult = (hi / lo) ** (1 / steps)5 lr, history = lo, []67 for i, batch in enumerate(loader):8 if i >= steps:9 break10 for g in opt.param_groups:11 g["lr"] = lr1213 loss = model(**batch).loss14 loss.backward()15 torch.nn.utils.clip_grad_norm_(16 [p for p in model.parameters() if p.requires_grad], 1.0)17 opt.step(); opt.zero_grad()1819 history.append((lr, loss.item()))20 if loss.item() > 4 * history[0][1]: # diverged21 break22 lr *= mult2324 return historyRead the plot, not the minimum. Loss falls as the rate rises, bottoms out, then explodes. Pick roughly one order of magnitude below the explosion point — for a typical QLoRA run the curve bottoms near 5e-4 and blows up around 3e-3, giving 2e-4 as the working choice.
If your grad norm sits pinned at
max_grad_normon every logged step, the learning rate is too high and clipping is silently rescuing you. Lower the rate rather than raising the clip.Rank and alpha
Rank sets how many independent directions the adapter can express, and it scales parameters linearly. For a Llama-2-7B with adapters on all seven linear modules per layer, each layer contributes r×78,080 parameters, so across 32 layers:
r Trainable params % of 6.74B Adapter size (BF16) Fits 8 19,988,480 0.297% 40 MB Style, tone, format 16 39,976,960 0.593% 80 MB Most instruction tuning 32 79,953,920 1.187% 160 MB Larger datasets, harder tasks 64 159,907,840 2.373% 320 MB Substantial behaviour change 128 319,815,680 4.746% 640 MB Rarely justified Match rank to data volume, not to ambition. A useful starting point: fewer than 1,000 examples, r=8; 1,000-10,000, r=16; 10,000-100,000, r=32; beyond that, r=64 and consider whether full fine-tuning is now affordable. Rank 128 on 2,000 examples is 320 million free parameters fitting 2,000 targets, and it overfits exactly as you would expect.
Alpha divides by rank in the forward pass — the adapter contributes (α/r)BAx — so it controls how loudly the adapter speaks, independent of how much it can say. Keeping α=2r fixes the scale at 2.0 for every rank, which is what makes a rank sweep interpretable: you are varying capacity while holding influence constant.
Python1for r in [8, 16, 32, 64]:2 cfg = LoraConfig(r=r, lora_alpha=2 * r, lora_dropout=0.05,3 target_modules=SEVEN_MODULES, task_type="CAUSAL_LM")4 # ... train, record best eval_loss, plot eval_loss against rIf eval loss improves from r=8 to r=16 and then flattens, take r=16 — the extra capacity is not being used. If it keeps improving to r=64, your task needs more capacity than a low-rank update naturally provides, and that is meaningful information about the task.
Before raising rank, try adding modules. Rank 8 across all seven linear projections (20.0M parameters) usually beats rank 32 on query and value alone (16.8M), for about the same parameter budget.
Checkpoints
Checkpoints are not only crash insurance. They are also how you recover the best model rather than the last one — and those are usually different, because eval loss typically bottoms out before training ends.
Python1args = TrainingArguments(2 output_dir="runs/support",3 save_strategy="steps",4 save_steps=50,5 save_total_limit=3, # keep 3 most recent; older are deleted6 eval_strategy="steps",7 eval_steps=50, # save_steps must be a multiple of this8 load_best_model_at_end=True,9 metric_for_best_model="eval_loss",10 greater_is_better=False,11)1213trainer = Trainer(..., args=args,14 callbacks=[EarlyStoppingCallback(early_stopping_patience=3)])Three details cause most checkpoint problems. With
load_best_model_at_end, the evaluation and save strategies must match andsave_stepsmust be a whole multiple ofeval_steps— keeping them equal is simplest. Get this wrong andTrainingArgumentsrefuses to start, which is annoying but far better than a silent failure.save_total_limitdeletes older checkpoints — but the best one is protected whenload_best_model_at_endis set, which is the reason to always set it. Andgreater_is_bettermust beFalsefor loss andTruefor accuracy or F1. It defaults correctly when the metric name ends inloss, but a custom metric where lower is better, such as an error rate, needs it set by hand; getting it backwards means the trainer faithfully restores your worst model.What a checkpoint costs
A full checkpoint stores model state, optimiser state and scheduler state so the run can resume exactly. For LoRA at r=16 across all modules, 39,976,960 trainable parameters:
Textadapter weights (FP32) 160 MBAdamW moments (2 x FP32) 320 MBscheduler + RNG + trainer state ~1 MB---------------------------------------per checkpoint ~481 MBx save_total_limit=3 ~1.44 GBFull fine-tuning, same model (no gradients saved):weights + moments + master ~94 GB per checkpointx 3 ~283 GBWith 8-bit optimiser states (
optim="paged_adamw_8bit") the moments shrink to 80 MB and a checkpoint costs about 241 MB. Either way, keeping three is trivially affordable — which is why turning checkpointing off to save disk is never the right trade.Resuming is one argument:
Pythontrainer.train(resume_from_checkpoint="runs/support/checkpoint-450")This restores optimiser moments, scheduler position and data ordering. Loading only the adapter weights and starting a fresh run is not the same thing — you lose the momentum state and restart the learning-rate schedule from warmup, which produces a visible discontinuity in the loss.
Sweeping automatically
Grid search over four hyperparameters at four values each is 256 runs. Bayesian search reaches a comparable result in 20-40, because it uses the results it already has to decide what to try next.
Python1import optuna23def objective(trial):4 lr = trial.suggest_float("lr", 5e-5, 5e-4, log=True)5 r = trial.suggest_categorical("r", [8, 16, 32, 64])6 drop = trial.suggest_float("dropout", 0.0, 0.15)78 model, tok = build_model(r=r, alpha=2 * r, dropout=drop)9 trainer = build_trainer(model, tok, lr=lr, max_steps=300)1011 trainer.train()12 metrics = trainer.evaluate()1314 trial.report(metrics["eval_loss"], step=300)15 if trial.should_prune(): # abandon hopeless trials early16 raise optuna.TrialPruned()17 return metrics["eval_loss"]1819study = optuna.create_study(20 direction="minimize",21 sampler=optuna.samplers.TPESampler(seed=42),22 pruner=optuna.pruners.MedianPruner(n_warmup_steps=100),23)24study.optimize(objective, n_trials=25)25print(study.best_params, study.best_value)Two things make this practical rather than ruinous. Cap each trial with
max_steps— a 300-step run ranks configurations almost as well as a 3,000-step run, at a tenth of the cost. And use a pruner, which kills trials performing below the median at a checkpoint; in practice it terminates around half the trials early.Sweep on
eval_lossonly if it correlates with what you care about. It often does not: two adapters with identical eval loss can differ substantially in schema compliance or refusal behaviour. If you have a task metric, optimise that instead.Tracking
Python1import wandb23wandb.init(project="support-finetune", name="r16-lr2e4",4 config={"r": 16, "alpha": 32, "lr": 2e-4,5 "effective_batch": 16, "max_length": 1024,6 "dataset_hash": "a41f9c", "git_sha": "3d9e1b2"})78args = TrainingArguments(..., report_to="wandb", run_name="r16-lr2e4",9 logging_steps=10)Log gradient norm, learning rate, and both losses. Gradient norm in particular is the earliest warning you get: a spike ten steps before the loss moves tells you a bad batch is coming through, and a norm pinned at the clip value tells you the rate is too high. Record the dataset hash and git commit in the config — six weeks later, "which data produced this?" is the first question and it should be a lookup.
Debugging playbook
Symptom Most likely cause Fix, in order Loss becomes NaN FP16 overflow, or LR far too high Switch to BF16; drop LR 10×; set max_grad_norm=0.3; look for an empty or corrupt example Loss rises steadily LR too high Divide LR by 5 and rerun 100 steps Loss flat from step 0 Nothing trainable, or all labels -100 Print trainable params; print one example's labels Loss falls to under 0.1 Memorising a small dataset Check the eval gap; reduce epochs; lower r; add data Train loss falls, eval loss rises Overfitting Early stopping; lora_dropout to 0.1; fewer epochs; lower r Both losses plateau high Underfitting Raise LR; raise r; add modules; train longer Loss good, generations bad Data formatting, not hyperparameters Check EOS token, prompt masking, inference template Grad norm pinned at the clip value LR too high Reduce LR by 3× CUDA OOM mid-run, not at start A long example in a later batch Cap max_length; sort by length; use paged optimiser Each epoch shows a loss step-change Data not shuffled Verify shuffling; check for ordered classes One diagnostic is worth more than the rest combined: overfit a batch of eight examples deliberately. Train on the same eight for 200 steps with everything else unchanged. Loss must reach near zero. If it does not, the problem is structural — masked labels, detached gradients, unattached adapters — and no hyperparameter will fix it. If it does, your pipeline is sound and the problem is a hyperparameter or the data. That test takes two minutes and cleanly partitions the space of causes.
What this means when you launch a run
Sequence your effort. Get one run to complete with default settings and a sane checkpoint policy before tuning anything — a mediocre model that exists beats a hypothetically excellent one that crashed at step 1,840. Then run the LR range test, which is 100 steps and eliminates the single largest source of failure. Then sweep rank over three values with alpha tied to 2r. Only after those three steps is anything else worth your GPU hours.
And set the checkpoint policy on the very first run, not once you have been burned.
save_steps=50,save_total_limit=3,load_best_model_at_end=True,metric_for_best_model="eval_loss",greater_is_better=False. Five lines, under two gigabytes of disk, and they convert an eviction from a lost day into a lost ten minutes.