Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

What is NF4 (NormalFloat-4) in QLoRA, and what problem was it designed to solve?


The positive half of the NF4 codebook0.0000.0800.1610.2460.3380.4410.5630.7231.000012345678exact zero0.24rounds hereabsmaxAn even 4-bit grid has steps of 0.133 everywhere; NF4 steps near zero are about 0.08.
Levels are packed where the bell curve puts most weights, so the common small values lose the least precision.

What you need to know

The problem with an evenly spaced grid

A plain 4-bit grid spreads its 16 values evenly from -1 to 1, about 0.133 apart. But most pretrained weights sit close to zero, in a bell curve. An even grid wastes many levels in the tails, where few weights live, and has poor resolution near zero, where most weights are.

NF4's answer

NF4 places its 16 values at the quantiles of a standard normal distribution, so each value covers about the same share of weights. The values used by bitsandbytes, rounded:

Text
-1.000 -0.696 -0.525 -0.395 -0.284 -0.185 -0.091  0.000 0.080  0.161  0.246  0.338  0.441  0.563  0.723  1.000

The steps near zero are about 0.08 apart; in the tails they are about 0.28 apart. It includes an exact zero, and has 7 negative and 8 positive values.

Blocks and scales

  1. Split the weights into blocks of 64.
  2. Scale — find the largest absolute value in the block (its "absmax") and divide the block by it, so values lie between -1 and 1.
  3. Round each value to the nearest NF4 level and store its 4-bit index.
  4. Dequantize when needed: look up the level and multiply by the block's absmax.

A worked example: a block has absmax 0.05, and one weight is 0.012. Normalised, that is 0.24. The nearest NF4 level is 0.246, so the stored value becomes 0.246 × 0.05 = 0.0123 — an error of 0.0003. On the even grid, the nearest level to 0.24 is 0.2, giving 0.010 — an error of 0.002, about seven times larger.

Blocks keep one large outlier weight from ruining the scale for the whole matrix. The per-block absmax values are what double quantization compresses further, from 32-bit to 8-bit.

What NF4 is and is not

It is a storage format. Every matrix multiply dequantizes the weights to bf16 first, so NF4 saves memory, not time — QLoRA is somewhat slower per step than bf16 LoRA. The QLoRA paper reported that NF4 with double quantization matched 16-bit fine-tuning quality on its benchmarks, and beat plain 4-bit float (FP4).

For serving, other formats are usually used: AWQ or GPTQ 4-bit integers, FP8, or newer hardware formats such as MXFP4 and NVFP4 on recent GPUs. These have faster inference kernels than bitsandbytes NF4.

A real-life example

A hospital's data rules say patient reports cannot leave its building. The IT team has one workstation with a 24 GB GPU and wants to fine-tune an 8B model as a medical-report summariser.

In bf16, the 8B base takes about 16 GB, leaving too little for activations with long reports. In NF4 with double quantization, the base takes about 5.5 GB, and training fits with 3,000-token reports. Before training, the team runs a quick check of the base model's perplexity on 500 held-out reports: 3.10 in bf16, 3.18 in NF4, and 3.41 in FP4. NF4 loses little, so they train with it. After training, they merge the adapter into a bf16 copy of the base and quantize that with AWQ for serving. (Made-up numbers for illustration.)

Follow-up questions to expect

  • "Why quantize in blocks instead of per tensor?" — One large outlier would stretch the scale for the whole tensor, crushing all other weights into a few levels. Blocks of 64 keep outliers local.
  • "Is NF4 used for activations too?" — No. It assumes a normal distribution, which fits weights. Activations have large outliers and are kept in bf16.
  • "NF4 versus GPTQ or AWQ?" — NF4 needs no calibration data and quantizes instantly at load time, which suits training. GPTQ and AWQ use calibration data to reduce error and have faster inference kernels, which suits serving.