Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

Compare FP32, FP16, BF16, and TF32. Why does precision choice matter during training?


How each format spends its bits8233.4e388103.4e3851065,504873.4e38Exponent bitsMantissa bitsLargest valueFP32TF32FP16BF16In BF16, 1.0 plus 0.001 rounds back to 1.0.
FP16 fails on range and BF16 is coarse on precision — training survives coarse steps but not overflow, so BF16 wins.

What you need to know

FormatBitsExponentMantissaLargest valueNote
FP3232823about 3.4 × 10^38Safe baseline; 4 bytes
TF3219 used, stored in 32810same as FP32NVIDIA tensor-core mode for FP32 matmuls
FP161651065,504Narrow range; needs loss scaling
BF161687about 3.4 × 10^38FP32 range, coarse steps

Why range matters more than precision for training

Gradients span a huge range. In FP16, anything smaller than about 6 × 10^-8 becomes zero, and anything above 65,504 becomes infinity. So FP16 training multiplies the loss by a large factor before backpropagation (loss scaling) and divides it back afterwards, skipping steps that overflow. BF16's range matches FP32, so this problem disappears. BF16's coarse steps are tolerable because training averages over many updates.

Where BF16's coarse steps do hurt

BF16 has only 7 mantissa bits: the next number after 1.0 is 1.0078. So 1.0 + 0.001 in BF16 is exactly 1.0 — the update vanishes. That is why mixed-precision training keeps fp32 master weights and optimizer states, and computes softmax, layer norm and loss in FP32.

TF32 is not on by default in PyTorch

TF32 speeds up FP32 matrix multiplies on Ampere and newer GPUs. Since PyTorch 1.12 it is off by default for matrix multiplies (cuDNN convolutions still use it by default). Turn it on for matmuls with torch.set_float32_matmul_precision("high"), or tf32=True in Hugging Face TrainingArguments.

Settings in practice

Python
from transformers import TrainingArgumentsargs = TrainingArguments(output_dir="out", bf16=True, tf32=True)  # Ampere or newer# older GPUs without bf16 (V100, T4): fp16=True, which adds loss scaling

Beyond these: FP8 is used for large-scale training on Hopper and Blackwell GPUs (DeepSeek-V3 was trained with FP8 mixed precision), with careful per-block scaling. For fine-tuning, BF16 remains the safe default.

A real-life example

A student team fine-tunes a 1.5B model for Hindi support on free T4 GPUs, which do not support BF16. They use fp16=True. Around step 800 the loss becomes NaN: some attention values exceeded 65,504 and turned into infinity, and loss scaling cannot prevent overflow in the forward pass.

They lower the learning rate and add gradient clipping, and the run survives — but it is fragile. When they move to an L4 GPU, which supports BF16, the same settings train smoothly with bf16=True and no loss scaling. Their interview answer later: "FP16 failed on range, not precision."

Follow-up questions to expect

  • "Why not train entirely in BF16, including the optimizer?" — Small updates get rounded away, as the 1.0 + 0.001 example shows. Keep master weights and optimizer states in FP32, or use optimizers designed for low-precision states.
  • "Is BF16 inference worse than FP32?" — For LLMs, the difference is usually negligible, which is why models ship in BF16.
  • "What is FP8?" — 8-bit floating point in two variants: E4M3 (more precision, used for weights and activations) and E5M2 (more range, used for gradients). It needs scaling factors and is mostly used in large training runs.