Fine-Tuning LLMs with LoRA, QLoRA and PEFT

Quantisation — FP32, FP16, BF16, INT8 and INT4


LoRA takes the optimiser state for a 7B model from 108 GB down to about 0.6 GB. Excellent. But the frozen base weights still sit there in BF16 at 2 bytes each: 6,738,415,616 × 2 = 13.5 GB. Add activations for a modest batch and you are at 18-20 GB. A 16 GB T4 cannot run it. A free Colab session cannot run it. And a 70B model at BF16 needs 140 GB before you have trained anything at all.

Here is the thing about those frozen weights: they are never updated. Not once. They are read, multiplied, and discarded. So why store them at full precision at all? If you could store each one in 4 bits instead of 16, the 13.5 GB becomes 3.4 GB and the whole problem changes shape.

That question is what quantisation answers, and QLoRA is the specific answer that made fine-tuning a 65B model on one 48 GB GPU possible.

What each precision costs on 6.74B weights4 bytes27.0 GBreference2 bytes13.5 GBno loss in practice2 bytes13.5 GBwider range,less mantissa1 byte6.7 GBsmall, needs outlier care0.5 bytes3.5 GBneeds NF4 and blocksPer weightBase modelQuality costFP32FP16BF16INT8NF4 (4-bit)13.5 GB plus activations will not fit a 16 GB T4; 3.5 GB plus activations will.
Halving precision halves memory linearly, but quality falls off a cliff between 8 and 4 bits unless the 4-bit format is chosen to match the weight distribution.

What quantisation is

Think of a kitchen scale. A laboratory balance reads to 0.001 g and can display any mass from 0 to 5,000 g — enormous range, fine resolution, and it costs a fortune. A cheap kitchen scale reads in 1 g steps up to 5 kg. For weighing flour, the cheap scale is fine. You lose the ability to distinguish 250.3 g from 250.7 g, and it does not matter.

Quantisation is choosing the cheap scale for numbers that do not need the expensive one. Formally: mapping values from a high-precision representation to a smaller set of discrete levels, storing the level index instead of the value, and reconstructing an approximation when you need it.

The two properties that matter are range (largest and smallest representable magnitude) and resolution (how finely you can distinguish nearby values). Every format trades them against each other within a fixed bit budget.

The floating-point formats

A floating-point number splits its bits into a sign, an exponent and a mantissa. The exponent sets the range; the mantissa sets the precision.

FormatBitsSign / Exp / MantissaMax magnitudeSmallest normalDecimal digits
FP32321 / 8 / 23≈3.4×1038\approx 3.4 \times 10^{38}≈1.2×10−38\approx 1.2 \times 10^{-38}~7
FP16161 / 5 / 1065,504≈6.1×10−5\approx 6.1 \times 10^{-5}~3
BF16161 / 8 / 7≈3.4×1038\approx 3.4 \times 10^{38}≈1.2×10−38\approx 1.2 \times 10^{-38}~2

Look at what BF16 did. It kept FP32's entire 8-bit exponent and paid for it by cutting the mantissa from 23 bits to 7. So BF16 has the same dynamic range as FP32 but coarser resolution, while FP16 has finer resolution than BF16 but a range that tops out at 65,504.

That range difference is the reason BF16 dominates modern training. During backpropagation, gradients routinely reach 10−810^{-8} and occasionally spike above 10510^{5}. In FP16, small gradients underflow to zero (below 6.1×10−56.1 \times 10^{-5} you fall into subnormals and then to zero) and large intermediate products overflow to infinity, which turns into NaN loss within a few steps. FP16 training therefore needs loss scaling — multiply the loss by 2162^{16} before backward, divide the gradients afterwards — an extra mechanism that occasionally fails. BF16 needs none of it.

FP16 and BF16 both cost two bytes. FP16 spends them on precision, BF16 on range. Training cares about range, which is why BF16 wins wherever the hardware supports it.

The catch: BF16 requires Ampere-generation hardware or newer (A100, A10, RTX 30-series and later). On a T4 or V100 you must use FP16, with loss scaling.

Integer quantisation, worked through

Integers have no exponent. An INT8 value is one of 256 evenly spaced levels; INT4 is one of 16. To represent real-valued weights with them you need a scale factor that maps the integer grid onto the range your weights actually occupy.

The scale factor

For symmetric quantisation to bb bits over a block of weights:

s=max⁡i∣wi∣2b−1−1,qi=round ⁣(wis),w^i=qi⋅ss = \frac{\max_i |w_i|}{2^{b-1} - 1}, \qquad q_i = \text{round}\!\left(\frac{w_i}{s}\right), \qquad \hat{w}_i = q_i \cdot s

You store the integers qiq_i (b bits each) plus one scale ss per block (usually FP16 or FP32). Reconstruction is a single multiply.

A real block of numbers

Take eight weights from a transformer layer and quantise to INT8, where 27−1=1272^7 - 1 = 127:

Text
weights:  0.0032  -0.0451   0.1203  -0.0087   0.0664  -0.1150   0.0021   0.0900absmax = 0.1203s      = 0.1203 / 127 = 0.000947w         w/s        q (rounded)   dequantised    error 0.0032     3.38            3        0.00284      0.00036-0.0451   -47.62          -48       -0.04547      0.00037 0.1203   127.00          127        0.12030      0.00000-0.0087    -9.19           -9       -0.00852      0.00018 0.0664    70.10           70        0.06631      0.00009-0.1150  -121.44         -121       -0.11462      0.00038 0.0021     2.22            2        0.00189      0.00021 0.0900    95.04           95        0.08999      0.00001max error 0.00038, and the quantisation step is 0.000947,so no error exceeds half a step (0.00047). That is the guarantee.

Eight FP32 numbers cost 32 bytes. Eight INT8 numbers plus one FP32 scale cost 8 + 4 = 12 bytes. And the reconstruction is accurate to about 0.3% of the block's largest value.

The outlier problem, and why block size matters

Now put a single large weight into that block — say 2.5, which happens in real transformers, particularly in certain feed-forward dimensions:

Text
weights:  0.0032  -0.0451   0.1203  -0.0087   2.5000  -0.1150   0.0021   0.0900absmax = 2.5s      = 2.5 / 127 = 0.01969 0.0032 / 0.01969 =  0.16  ->  q = 0   ->  dequantised 0.0000   (the value is gone)-0.0451 / 0.01969 = -2.29  ->  q = -2  ->  dequantised -0.0394  (13% error) 0.1203 / 0.01969 =  6.11  ->  q = 6   ->  dequantised 0.1181 2.5000 / 0.01969 = 127.0  ->  q = 127 ->  dequantised 2.5000   (perfect)

One outlier stretched the scale by a factor of 21 and destroyed every small weight in the block. This is exactly why naive whole-tensor quantisation of large language models fails badly, and why every working scheme uses block-wise quantisation: split the tensor into blocks of 64 or 128 weights, compute a separate scale for each. An outlier then damages only its own 64 neighbours rather than the whole matrix.

INT8 to INT4: the resolution cliff

INT4 has 16 levels, so symmetric quantisation uses 23−1=72^3 - 1 = 7:

Text
Same block, absmax 0.1203, no outlier.INT8:  s = 0.1203 / 127 = 0.000947   step 0.000947INT4:  s = 0.1203 /   7 = 0.017186   step 0.017186The INT4 step is 127/7 = 18.1x coarser. 0.0032 / 0.017186 = 0.19  ->  q = 0  ->  0.0000  (100% of the value lost) 0.0664 / 0.017186 = 3.86  ->  q = 4  ->  0.0687  (3.5% error)-0.1150 / 0.017186 = -6.69 ->  q = -7 ->  -0.1203 (4.6% error)

Small weights vanish entirely. Halving the bit width does not double the error; it makes the step, and with it the worst-case error, about 18 times larger. This is why plain INT4 with uniform levels degrades a language model noticeably — and it is the specific problem NF4 was designed to solve.

Post-training quantisation versus quantisation-aware training

Post-training quantisation (PTQ)Quantisation-aware training (QAT)
When it happensAfter training, on a finished modelDuring training
How it worksCompute scales from the weights; optionally calibrate activation ranges on 128-512 sample inputsSimulate quantisation in the forward pass; use a straight-through estimator so gradients flow
CostMinutes; no training data needed for weight-only schemesA full training run
Accuracy at 8 bitsUsually within 1%Essentially lossless
Accuracy at 4 bits and belowNoticeable degradation with uniform gridsBest available
When to useDeployment of an existing checkpoint; the default choiceVery low bit widths where PTQ is not good enough and you can afford to retrain

QLoRA is neither, exactly, and this is worth being precise about. The base weights are quantised post-hoc to 4 bits and frozen there permanently — pure PTQ, never updated, never fine-tuned back. The trainable LoRA adapters live in BF16 alongside them. Gradients flow through the quantised base (it must be dequantised on the fly to compute the matmul) but never into it. So it is a quantised frozen backbone with full-precision learnable corrections — a third pattern, and the reason it recovers full 16-bit fine-tuning quality despite a 4-bit base.

The memory that precision actually buys

PrecisionBytes/param7B model13B model70B model
FP32427.0 GB52.0 GB280 GB
FP16 / BF16213.5 GB26.0 GB140 GB
INT816.7 GB13.0 GB70 GB
NF4 / INT40.53.4 GB6.5 GB35 GB

Read the 70B row across. At BF16 you need two 80 GB A100s just to hold the weights. At NF4 the weights need about 35 GB, which fits on one 48 GB card (an A6000 or L40S) with room left over for adapters and activations. That is the scale of result the QLoRA paper demonstrated: a 65B model fine-tuned on a single 48 GB GPU. That single row is why quantisation matters more than any other memory technique.

Add the LoRA arithmetic and you get the full QLoRA budget for a 7B model at rank 16 across all modules:

Text
Linear layers in NF4 (6.48B params)     3.24 GBQuantisation constants (double quant)   0.10 GBEmbeddings + LM head (262M params,  not quantised; FP32 after prep)       1.05 GBAdapter params (40M, FP32)              0.16 GBAdapter gradients (FP32)                0.16 GBAdamW moments (2 x FP32)                0.32 GBActivations, batch 4 x 1024 tokens,  with gradient checkpointing          ~1.5  GB-----------------------------------------------Total                                  ~6.5  GBCompare: full fine-tuning, same model            ~125 GB         LoRA on a BF16 base, same model          ~16 GB

Notice what is not 4-bit. bitsandbytes quantises only the linear layers inside the transformer blocks. The token embeddings and the LM head (2 × 32,000 × 4,096 = 262 million parameters) stay in 16-bit, and prepare_model_for_kbit_training, covered below, upcasts them to FP32: 0.52 GB becomes 1.05 GB. For Llama-2 that is a small line. For current models with vocabularies around 150,000 tokens it is not: in Qwen3-8B the embeddings and LM head hold about 1.24 billion parameters, roughly 2.5 GB in BF16, which is why its 4-bit footprint is about 6 GB rather than 4.

QLoRA's three specific contributions

4-bit NormalFloat

INT4 spreads its 16 levels evenly across the range. But trained neural network weights are not uniformly distributed — they are approximately zero-centred and roughly normal, with most of the mass bunched near zero and a thin tail out to the extremes. Uniform levels therefore waste resolution on the sparse tails and starve the dense centre, which is exactly what destroyed the small weights in the worked example above.

NF4 places its 16 levels at the quantiles of a standard normal distribution, so that each level covers an equal share of probability mass rather than an equal share of the numeric range. Levels are packed tightly near zero, where the weights are, and spread widely in the tails, where they are not. Zero is represented exactly, which preserves genuine zeros and keeps the format symmetric.

Text
Uniform INT4 levels (normalised):  -1.00  -0.86  -0.71  -0.57  -0.43  -0.29  -0.14  0.00   0.14   0.29   0.43   0.57   0.71   0.86   1.00  -- evenly spaced; wide gaps everywhere near zeroNF4 levels (normalised, approximate):  -1.00  -0.70  -0.53  -0.39  -0.28  -0.18  -0.09  0.00   0.08   0.16   0.25   0.34   0.44   0.56   0.72   1.00  -- gap near zero is ~0.08 vs 0.14 for uniform

The claim in the QLoRA work is that NF4 is information-theoretically optimal for normally distributed data, and empirically it recovers close to BF16 fine-tuning quality where uniform INT4 does not.

Double quantisation

Block-wise quantisation needs one scale constant per block. With block size 64 and an FP32 constant, that overhead is:

32 bits64 params=0.5 bits per parameter\frac{32 \text{ bits}}{64 \text{ params}} = 0.5 \text{ bits per parameter}

Half a bit on top of four is a 12.5% overhead — not nothing. Double quantisation quantises the constants themselves: group 256 blocks together, quantise their FP32 scales to 8 bits, and keep one FP32 scale for that group of scales:

Text
First level:   8 bits per block-64 scale     ->  8 / 64        = 0.125   bits/paramSecond level:  32 bits per group of 256      -> 32 / (64x256)  = 0.00195 bits/param                                                 ------------------------                                                 total          0.127   bits/paramSaving: 0.500 - 0.127 = 0.373 bits per parameter

For a 65B model that is 65×109×0.373/865 \times 10^{9} \times 0.373 / 8 bytes ≈ 3.0 GB recovered, for free, with no measurable quality cost. For a 7B model it is about 0.3 GB: small, but free, and it counts on a card that is already nearly full.

Paged optimisers

Memory usage during training is not flat. It spikes — when a long sequence arrives in the batch, when gradient checkpointing recomputes a block, when an evaluation pass runs. A run that averages 14 GB on a 16 GB card can die on a single 17 GB spike four hours in.

Paged optimisers allocate the optimiser state in NVIDIA unified memory, which allows pages to be evicted to CPU RAM automatically when GPU memory is exhausted and paged back when accessed — the same idea as operating-system virtual memory. The spike gets absorbed by system RAM instead of crashing the run. It is slower when paging actually occurs, which is the point: slow beats dead.

NF4 makes 4 bits accurate enough to train through, double quantisation reclaims the overhead of making it block-wise, and paged optimisers stop transient spikes from killing an eight-hour run. Together they are what turns "4-bit is an interesting idea" into "4-bit is what people actually use".

Configuring it

Python
import torchfrom transformers import AutoModelForCausalLM, BitsAndBytesConfigbnb_config = BitsAndBytesConfig(    load_in_4bit=True,    bnb_4bit_quant_type="nf4",            # "nf4" or "fp4"; nf4 is better for weights    bnb_4bit_use_double_quant=True,       # reclaim the 0.373 bits/param    bnb_4bit_compute_dtype=torch.bfloat16, # dtype the matmuls run in)model = AutoModelForCausalLM.from_pretrained(    "meta-llama/Llama-2-7b-hf",    quantization_config=bnb_config,    device_map="auto",)print(model.get_memory_footprint() / 1e9, "GB")   # ~3.9 GB: 4-bit layers + 16-bit embeddings

The parameter people misconfigure is bnb_4bit_compute_dtype. Storage is 4-bit; computation is not. Every matmul dequantises the relevant block into this dtype, multiplies, and discards the dequantised copy. Leave it at the default, which is still FP32 in current transformers, and you get correct results at roughly half the speed of BF16 for no benefit. Set it to BF16 on Ampere or newer; FP16 on older cards.

Note also that quantised models must be prepared before attaching adapters:

Python
from peft import prepare_model_for_kbit_training, LoraConfig, get_peft_modelmodel = prepare_model_for_kbit_training(model, use_gradient_checkpointing=True)# casts every non-quantised weight (layer norms, embeddings, LM head)# to FP32, enables input gradients so activations can flow back# through the frozen quantised layers, and turns on checkpointingmodel = get_peft_model(model, LoraConfig(r=16, lora_alpha=32,                                         target_modules=["q_proj", "v_proj"],                                         task_type="CAUSAL_LM"))

Pitfalls that cost people a day each

Trying to full-fine-tune a 4-bit model. A 4-bit weight has 16 possible values. A gradient step of 10−510^{-5} rounds to exactly the same value. Quantised weights are frozen by construction — you must attach trainable high-precision adapters, which is why QLoRA is LoRA on a quantised base and not something simpler.

Expecting 4-bit to be faster. Quantisation optimises memory, not necessarily speed. Every matmul pays a dequantisation cost, so a 4-bit forward pass is often 10-30% slower per token than BF16 on hardware that could hold the BF16 model. The win is that it fits at all, and that a larger batch becomes possible.

Skipping prepare_model_for_kbit_training. Without it, gradients do not propagate through the frozen quantised layers to reach adapters in earlier blocks, and layer norms in low precision destabilise the loss. The symptom is a loss that barely moves or goes to NaN.

Using FP16 compute dtype on a card without loss scaling configured. The classic symptom is loss dropping normally for 200 steps then jumping to NaN and staying there. Switch to BF16 if the hardware allows.

Quantising and then evaluating only perplexity. A quantised model can hold nearly the same perplexity while losing measurably on structured tasks — long-context retrieval and exact-format output degrade first. Evaluate on the thing you actually do.

Merging a LoRA adapter into a 4-bit base. The result is a 4-bit model whose weights were adjusted then re-rounded to 16 levels, and quality drops. Correct procedure: reload the base in BF16, apply the adapter to that, then merge. The adapter was trained against a 4-bit forward pass but transfers fine.

What this means when you set up a run

Work out your budget in bytes before you launch anything. Parameter count × bytes-per-parameter gives the frozen base; add roughly 16 bytes per trainable parameter for gradients and optimiser state; add 1-3 GB for activations if gradient checkpointing is on, ten times that if it is off. Compare against your GPU. That calculation takes two minutes and saves the twenty-minute cycle of launching a job, waiting for the model to load, and reading an OOM trace.

Then pick the precision from the constraint rather than from habit. If BF16 fits, use BF16 — it is faster and has no quantisation error at all. If it does not fit, go to NF4 with double quantisation and BF16 compute, and accept a small speed penalty in exchange for a run that exists. There is no reason to reach for INT8 for LoRA fine-tuning in between; NF4 with the QLoRA machinery is both smaller and, for this use, better behaved.