Fine-Tuning LLMs with LoRA, QLoRA and PEFT

Hyperparameters and Checkpoints — Tuning a Run


A run has been going for six hours. Loss started at 2.1, fell steadily to 0.91, and at step 1,840 the pod is evicted. There are no checkpoints, because save_strategy was left at its default of "steps" with save_steps=500 and someone changed it to "no" a week ago to save disk. Six hours of A100 time, gone, with nothing to show for it — not even the loss history, which lived in the process.

A second team avoids that trap but hits a different one. Their loss curve looks fine, and the finished model is worse than the base model at everything. The cause: learning_rate=2e-3, ten times too high. The loss decreased because the model learned to reproduce the training set's surface patterns while its general capabilities collapsed. The curve told them nothing.

Hyperparameters decide whether a run learns the right thing; checkpoints decide whether you keep it. Both are cheap to get right and expensive to get wrong.

Picking rank, and what each rank buys0.5 M4 MB8narrow format fixes1.1 M9 MB16most task adaptation4.2 M16 MB32the usual default8.4 M33 MB64domain vocabulary16.8 M65 MB128nearfull-tune capacityTrainableAdapter fileTypical alphaWhere it earns itr = 4r = 8r = 16r = 32r = 64Alpha near twice the rank keeps the effective update scale roughly constant as r moves.
Rank buys capacity, not quality — past the point where the task's update actually is low-rank, higher r only overfits faster and costs more disk.

What you are actually tuning

ParameterTypical range (QLoRA, 7B)ImpactTune it?
learning_rate5e-5 to 5e-4Very highAlways — first thing
lora r8 to 64HighUsually — sweep 3 values
lora_alpha2r as a defaultMediumTie to r; sweep rarely
num_train_epochs1 to 5HighUse early stopping instead
effective batch size8 to 128MediumSet by memory, then match LR
warmup_steps (as a fraction)0.03 to 0.10Low-mediumRarely — 0.03 works
lora_dropout0.0 to 0.1Low-mediumRaise if overfitting
weight_decay0.0 to 0.1Low0.01 and move on
max_grad_norm0.3 to 1.0Low, until it is notLower it if you see spikes
target_modules2 or 7 modulesHighYes — often beats raising r

Spend your sweep budget on the top of that table. Learning rate and rank between them explain most of the variance between a good and a bad run; weight decay explains almost none.

Effective batch size

Three separate settings combine into the number that actually matters:

effective batch=per_device_batch×grad_accum_steps×num_gpus\text{effective batch} = \text{per\_device\_batch} \times \text{grad\_accum\_steps} \times \text{num\_gpus}

Gradient accumulation runs several forward/backward passes, sums the gradients, and only then takes one optimiser step. Mathematically this is close to a single large batch, but peak memory is set by the per-device batch alone. That is the lever: you can train with an effective batch of 64 on a card that can only hold 4 sequences at a time.

Text
per_device=16, accum=1   ->  effective 16, memory for 16 sequences   (may OOM)per_device=4,  accum=4   ->  effective 16, memory for 4 sequences    (fits)per_device=1,  accum=16  ->  effective 16, memory for 1 sequence     (fits anywhere)

The cost is throughput. Sixteen separate forward passes have more per-step overhead than one batched pass, so per_device=1, accum=16 might run 30-40% slower per optimiser step than per_device=4, accum=4. Use the largest per-device batch that fits, then make up the rest with accumulation.

What batch size does to learning

  • Small effective batch (4-8): noisy gradients. The noise acts as a regulariser and can help on small datasets, but training is unstable and one bad example can move the weights a long way.
  • Medium (16-32): the sweet spot for most instruction fine-tuning. Enough averaging to be stable, enough noise to generalise.
  • Large (64-256): smooth, stable, fast in wall-clock terms if you have the memory — but each step sees the same data averaged more heavily, so you need proportionally fewer steps and often a higher learning rate to compensate.

The rough correspondence is the square-root rule: if you multiply batch size by 4, multiply learning rate by about 2. It is a heuristic, not a law, but it stops the common mistake of quadrupling the batch to fix an OOM and then wondering why the model underfits.

GPUMemoryModel / precisionSuggested per_device × accumMax length
T416 GB7B, NF4, fp161 × 16512
RTX 409024 GB7B, NF4, bf164 × 41024
A10G24 GB7B, NF4, bf164 × 41024
A10040 GB13B, NF4, bf164 × 42048
A10080 GB70B, NF4, bf162 × 82048

Learning rate

Why LoRA wants a bigger one than you expect

Full fine-tuning of a 7B model typically uses 1e-5 to 5e-5. LoRA typically uses 1e-4 to 3e-4 — five to ten times higher. The reason is that LoRA's matrix BB starts at exactly zero, so the adapter contributes nothing at step 0 and has to travel a meaningful distance before it does anything at all. There is also no risk of wrecking the pretrained weights, because they are frozen. A learning rate that would destroy a full fine-tune is merely brisk for an adapter.

Schedules

A constant learning rate is wrong at both ends of training. At the start, the adapter is random and the first few gradients are large and unreliable; a full-size step here can knock the run into a bad region it never recovers from. At the end, you want small refinements, and a full-size step overshoots the minimum repeatedly.

Warmup plus cosine decay fixes both:

During warmup, for step t≤Twt \le T_w:

η(t)=ηmax⁡⋅tTw\eta(t) = \eta_{\max} \cdot \frac{t}{T_w}

After warmup, with TT total steps:

η(t)=ηmax⁡2(1+cos⁡(π⋅t−TwT−Tw))\eta(t) = \frac{\eta_{\max}}{2}\left(1 + \cos\left(\pi \cdot \frac{t - T_w}{T - T_w}\right)\right)

Worked through, with ηmax⁡=2×10−4\eta_{\max} = 2 \times 10^{-4}, T=375T = 375 total steps and Tw=30T_w = 30 warmup steps:

StepProgress through decaycos termLearning rate
0——0
15half of warmup—1.00e-4
30warmup complete1.0002.00e-4
10070/345 = 0.2030.8031.80e-4
200170/345 = 0.4930.0231.02e-4
300270/345 = 0.783-0.7762.24e-5
3751.000-1.0000

Notice the shape: the rate stays near its maximum for the first third, then falls away sharply. That is deliberate — most of the learning happens early, and the long tail of small steps is what makes the final model stable.

ScheduleBehaviourUse when
constantFlat throughoutDebugging only
constant_with_warmupRamp then flatContinuing a run; unknown total length
linearRamp then straight decline to 0Simple, predictable, fine
cosineRamp then smooth decline to 0Default for fine-tuning
cosine_with_restartsRepeated decay cyclesLong runs escaping plateaus

Finding the rate

Do not guess and do not sweep blindly. Run a short LR range test: 100 steps, increasing the rate exponentially, and plot loss against rate.

Python
import math, torchdef lr_range_test(model, loader, opt, lo=1e-6, hi=1e-2, steps=100):    mult = (hi / lo) ** (1 / steps)    lr, history = lo, []    for i, batch in enumerate(loader):        if i >= steps:            break        for g in opt.param_groups:            g["lr"] = lr        loss = model(**batch).loss        loss.backward()        torch.nn.utils.clip_grad_norm_(            [p for p in model.parameters() if p.requires_grad], 1.0)        opt.step(); opt.zero_grad()        history.append((lr, loss.item()))        if loss.item() > 4 * history[0][1]:      # diverged            break        lr *= mult    return history

Read the plot, not the minimum. Loss falls as the rate rises, bottoms out, then explodes. Pick roughly one order of magnitude below the explosion point — for a typical QLoRA run the curve bottoms near 5e-4 and blows up around 3e-3, giving 2e-4 as the working choice.

If your grad norm sits pinned at max_grad_norm on every logged step, the learning rate is too high and clipping is silently rescuing you. Lower the rate rather than raising the clip.

Rank and alpha

Rank sets how many independent directions the adapter can express, and it scales parameters linearly. For a Llama-2-7B with adapters on all seven linear modules per layer, each layer contributes r×78,080r \times 78{,}080 parameters, so across 32 layers:

rTrainable params% of 6.74BAdapter size (BF16)Fits
819,988,4800.297%40 MBStyle, tone, format
1639,976,9600.593%80 MBMost instruction tuning
3279,953,9201.187%160 MBLarger datasets, harder tasks
64159,907,8402.373%320 MBSubstantial behaviour change
128319,815,6804.746%640 MBRarely justified

Match rank to data volume, not to ambition. A useful starting point: fewer than 1,000 examples, r=8r = 8; 1,000-10,000, r=16r = 16; 10,000-100,000, r=32r = 32; beyond that, r=64r = 64 and consider whether full fine-tuning is now affordable. Rank 128 on 2,000 examples is 320 million free parameters fitting 2,000 targets, and it overfits exactly as you would expect.

Alpha divides by rank in the forward pass — the adapter contributes (α/r)BAx(\alpha/r) BAx — so it controls how loudly the adapter speaks, independent of how much it can say. Keeping α=2r\alpha = 2r fixes the scale at 2.0 for every rank, which is what makes a rank sweep interpretable: you are varying capacity while holding influence constant.

Python
for r in [8, 16, 32, 64]:    cfg = LoraConfig(r=r, lora_alpha=2 * r, lora_dropout=0.05,                     target_modules=SEVEN_MODULES, task_type="CAUSAL_LM")    # ... train, record best eval_loss, plot eval_loss against r

If eval loss improves from r=8r=8 to r=16r=16 and then flattens, take r=16r=16 — the extra capacity is not being used. If it keeps improving to r=64r=64, your task needs more capacity than a low-rank update naturally provides, and that is meaningful information about the task.

Before raising rank, try adding modules. Rank 8 across all seven linear projections (20.0M parameters) usually beats rank 32 on query and value alone (16.8M), for about the same parameter budget.

Checkpoints

Checkpoints are not only crash insurance. They are also how you recover the best model rather than the last one — and those are usually different, because eval loss typically bottoms out before training ends.

Python
args = TrainingArguments(    output_dir="runs/support",    save_strategy="steps",    save_steps=50,    save_total_limit=3,               # keep 3 most recent; older are deleted    eval_strategy="steps",    eval_steps=50,                    # save_steps must be a multiple of this    load_best_model_at_end=True,    metric_for_best_model="eval_loss",    greater_is_better=False,)trainer = Trainer(..., args=args,                  callbacks=[EarlyStoppingCallback(early_stopping_patience=3)])

Three details cause most checkpoint problems. With load_best_model_at_end, the evaluation and save strategies must match and save_steps must be a whole multiple of eval_steps — keeping them equal is simplest. Get this wrong and TrainingArguments refuses to start, which is annoying but far better than a silent failure. save_total_limit deletes older checkpoints — but the best one is protected when load_best_model_at_end is set, which is the reason to always set it. And greater_is_better must be False for loss and True for accuracy or F1. It defaults correctly when the metric name ends in loss, but a custom metric where lower is better, such as an error rate, needs it set by hand; getting it backwards means the trainer faithfully restores your worst model.

What a checkpoint costs

A full checkpoint stores model state, optimiser state and scheduler state so the run can resume exactly. For LoRA at r=16r=16 across all modules, 39,976,960 trainable parameters:

Text
adapter weights (FP32)          160 MBAdamW moments (2 x FP32)        320 MBscheduler + RNG + trainer state   ~1 MB---------------------------------------per checkpoint                 ~481 MBx save_total_limit=3          ~1.44 GBFull fine-tuning, same model (no gradients saved):weights + moments + master     ~94 GB per checkpointx 3                            ~283 GB

With 8-bit optimiser states (optim="paged_adamw_8bit") the moments shrink to 80 MB and a checkpoint costs about 241 MB. Either way, keeping three is trivially affordable — which is why turning checkpointing off to save disk is never the right trade.

Resuming is one argument:

Python
trainer.train(resume_from_checkpoint="runs/support/checkpoint-450")

This restores optimiser moments, scheduler position and data ordering. Loading only the adapter weights and starting a fresh run is not the same thing — you lose the momentum state and restart the learning-rate schedule from warmup, which produces a visible discontinuity in the loss.

Sweeping automatically

Grid search over four hyperparameters at four values each is 256 runs. Bayesian search reaches a comparable result in 20-40, because it uses the results it already has to decide what to try next.

Python
import optunadef objective(trial):    lr   = trial.suggest_float("lr", 5e-5, 5e-4, log=True)    r    = trial.suggest_categorical("r", [8, 16, 32, 64])    drop = trial.suggest_float("dropout", 0.0, 0.15)    model, tok = build_model(r=r, alpha=2 * r, dropout=drop)    trainer = build_trainer(model, tok, lr=lr, max_steps=300)    trainer.train()    metrics = trainer.evaluate()    trial.report(metrics["eval_loss"], step=300)    if trial.should_prune():          # abandon hopeless trials early        raise optuna.TrialPruned()    return metrics["eval_loss"]study = optuna.create_study(    direction="minimize",    sampler=optuna.samplers.TPESampler(seed=42),    pruner=optuna.pruners.MedianPruner(n_warmup_steps=100),)study.optimize(objective, n_trials=25)print(study.best_params, study.best_value)

Two things make this practical rather than ruinous. Cap each trial with max_steps — a 300-step run ranks configurations almost as well as a 3,000-step run, at a tenth of the cost. And use a pruner, which kills trials performing below the median at a checkpoint; in practice it terminates around half the trials early.

Sweep on eval_loss only if it correlates with what you care about. It often does not: two adapters with identical eval loss can differ substantially in schema compliance or refusal behaviour. If you have a task metric, optimise that instead.

Tracking

Python
import wandbwandb.init(project="support-finetune", name="r16-lr2e4",           config={"r": 16, "alpha": 32, "lr": 2e-4,                   "effective_batch": 16, "max_length": 1024,                   "dataset_hash": "a41f9c", "git_sha": "3d9e1b2"})args = TrainingArguments(..., report_to="wandb", run_name="r16-lr2e4",                         logging_steps=10)

Log gradient norm, learning rate, and both losses. Gradient norm in particular is the earliest warning you get: a spike ten steps before the loss moves tells you a bad batch is coming through, and a norm pinned at the clip value tells you the rate is too high. Record the dataset hash and git commit in the config — six weeks later, "which data produced this?" is the first question and it should be a lookup.

Debugging playbook

SymptomMost likely causeFix, in order
Loss becomes NaNFP16 overflow, or LR far too highSwitch to BF16; drop LR 10×; set max_grad_norm=0.3; look for an empty or corrupt example
Loss rises steadilyLR too highDivide LR by 5 and rerun 100 steps
Loss flat from step 0Nothing trainable, or all labels -100Print trainable params; print one example's labels
Loss falls to under 0.1Memorising a small datasetCheck the eval gap; reduce epochs; lower r; add data
Train loss falls, eval loss risesOverfittingEarly stopping; lora_dropout to 0.1; fewer epochs; lower r
Both losses plateau highUnderfittingRaise LR; raise r; add modules; train longer
Loss good, generations badData formatting, not hyperparametersCheck EOS token, prompt masking, inference template
Grad norm pinned at the clip valueLR too highReduce LR by 3×
CUDA OOM mid-run, not at startA long example in a later batchCap max_length; sort by length; use paged optimiser
Each epoch shows a loss step-changeData not shuffledVerify shuffling; check for ordered classes

One diagnostic is worth more than the rest combined: overfit a batch of eight examples deliberately. Train on the same eight for 200 steps with everything else unchanged. Loss must reach near zero. If it does not, the problem is structural — masked labels, detached gradients, unattached adapters — and no hyperparameter will fix it. If it does, your pipeline is sound and the problem is a hyperparameter or the data. That test takes two minutes and cleanly partitions the space of causes.

What this means when you launch a run

Sequence your effort. Get one run to complete with default settings and a sane checkpoint policy before tuning anything — a mediocre model that exists beats a hypothetically excellent one that crashed at step 1,840. Then run the LR range test, which is 100 steps and eliminates the single largest source of failure. Then sweep rank over three values with alpha tied to 2r2r. Only after those three steps is anything else worth your GPU hours.

And set the checkpoint policy on the very first run, not once you have been burned. save_steps=50, save_total_limit=3, load_best_model_at_end=True, metric_for_best_model="eval_loss", greater_is_better=False. Five lines, under two gigabytes of disk, and they convert an eviction from a lost day into a lost ten minutes.