Course Content
Fine-Tuning LLMs
6 sections · 52 lessons
Compare FP32, FP16, BF16, and TF32. Why does precision choice matter during training?
What you need to know
| Format | Bits | Exponent | Mantissa | Largest value | Note |
|---|---|---|---|---|---|
| FP32 | 32 | 8 | 23 | about 3.4 × 10^38 | Safe baseline; 4 bytes |
| TF32 | 19 used, stored in 32 | 8 | 10 | same as FP32 | NVIDIA tensor-core mode for FP32 matmuls |
| FP16 | 16 | 5 | 10 | 65,504 | Narrow range; needs loss scaling |
| BF16 | 16 | 8 | 7 | about 3.4 × 10^38 | FP32 range, coarse steps |
Why range matters more than precision for training
Gradients span a huge range. In FP16, anything smaller than about 6 × 10^-8 becomes zero, and anything above 65,504 becomes infinity. So FP16 training multiplies the loss by a large factor before backpropagation (loss scaling) and divides it back afterwards, skipping steps that overflow. BF16's range matches FP32, so this problem disappears. BF16's coarse steps are tolerable because training averages over many updates.
Where BF16's coarse steps do hurt
BF16 has only 7 mantissa bits: the next number after 1.0 is 1.0078. So 1.0 + 0.001 in BF16 is exactly 1.0 — the update vanishes. That is why mixed-precision training keeps fp32 master weights and optimizer states, and computes softmax, layer norm and loss in FP32.
TF32 is not on by default in PyTorch
TF32 speeds up FP32 matrix multiplies on Ampere and newer GPUs. Since PyTorch 1.12 it is off by default for matrix multiplies (cuDNN convolutions still use it by default). Turn it on for matmuls with torch.set_float32_matmul_precision("high"), or tf32=True in Hugging Face TrainingArguments.
Settings in practice
1from transformers import TrainingArguments23args = TrainingArguments(output_dir="out", bf16=True, tf32=True) # Ampere or newer4# older GPUs without bf16 (V100, T4): fp16=True, which adds loss scalingBeyond these: FP8 is used for large-scale training on Hopper and Blackwell GPUs (DeepSeek-V3 was trained with FP8 mixed precision), with careful per-block scaling. For fine-tuning, BF16 remains the safe default.
A real-life example
A student team fine-tunes a 1.5B model for Hindi support on free T4 GPUs, which do not support BF16. They use fp16=True. Around step 800 the loss becomes NaN: some attention values exceeded 65,504 and turned into infinity, and loss scaling cannot prevent overflow in the forward pass.
They lower the learning rate and add gradient clipping, and the run survives — but it is fragile. When they move to an L4 GPU, which supports BF16, the same settings train smoothly with bf16=True and no loss scaling. Their interview answer later: "FP16 failed on range, not precision."
Follow-up questions to expect
- "Why not train entirely in BF16, including the optimizer?" — Small updates get rounded away, as the 1.0 + 0.001 example shows. Keep master weights and optimizer states in FP32, or use optimizers designed for low-precision states.
- "Is BF16 inference worse than FP32?" — For LLMs, the difference is usually negligible, which is why models ship in BF16.
- "What is FP8?" — 8-bit floating point in two variants: E4M3 (more precision, used for weights and activations) and E5M2 (more range, used for gradients). It needs scaling factors and is mostly used in large training runs.