Course Content
Fine-Tuning LLMs
6 sections · 52 lessons
What is mixed-precision training and what benefits/risks does it introduce?
What you need to know
The number formats themselves (FP32, FP16, BF16, TF32) are compared in Section 3. This lesson is about how training mixes them.
How one training step works
- Keep master weights in FP32 — the optimizer owns these.
- Autocast the forward pass — matrix multiplies run in bf16; PyTorch's autocast keeps softmax, layer norm and the loss in FP32.
- Backward pass — gradients flow back through the same mixed operations.
- Update in FP32 — the optimizer applies the gradients to the FP32 master weights.
Why the FP32 master copy matters
bf16 has only 8 bits of precision. Near 1.0, the gap between two bf16 numbers is 2^-7, about 0.0078. So 1.0 + 0.0001 in bf16 rounds straight back to 1.0 — a typical small update just disappears. In FP32 the update is kept. Over thousands of steps, those lost updates are the difference between a model that learns and one that stalls.
Benefits, with numbers
- Speed: an A100 is rated at 312 TFLOPS for dense BF16 on tensor cores, against 19.5 TFLOPS for plain FP32 (156 with TF32). Real speed-ups are smaller than this ratio but large.
- Activation memory: 16-bit activations take half the space, so longer sequences or bigger micro-batches fit.
What it does not save: for full fine-tuning with Adam, memory per parameter stays about 16 bytes (2 for bf16 weights, 2 for gradients, 4 for the FP32 master, 8 for Adam), which is the same as pure FP32 training.
How to turn it on
from trl import SFTConfigSFTConfig(output_dir="out", bf16=True) # Ampere, Ada, Hopper, Blackwell# SFTConfig(output_dir="out", fp16=True) # T4 or V100: adds a dynamic loss scalerWith LoRA, the frozen base is usually loaded directly in bf16 (dtype=torch.bfloat16) with no master copy, since it never updates. PEFT keeps the trainable adapter weights in FP32 by default, so the part that learns still gets precise updates.
Risks
- FP16 range: FP16 cannot represent very small gradients, so they become zero. Dynamic loss scaling multiplies the loss before the backward pass and divides the gradients back before the step. Occasional skipped steps are normal; constant NaNs mean a real instability.
- Rounding drift: sums and updates done in 16-bit lose small values. Keep them in FP32.
- Reproducibility: results vary slightly across GPUs and kernel choices.
FP8 training (on H100 and newer) goes further; DeepSeek-V3 was trained with FP8 mixed precision. For fine-tuning in 2026, bf16 is still the normal default.
A real-life example
A student team fine-tunes a brand-voice product-description writer on a free T4 GPU. The T4 has no bf16 support, so they use fp16=True. Around step 300 the loss becomes NaN.
The loss scaler is already on, so they look further: the gradient norm has been climbing for 100 steps, and the learning rate is 3e-4. They lower it to 1e-4, keep gradient clipping at 1.0, and the run finishes. Later they move to an L4 (Ada generation, bf16 supported), switch to bf16=True, and the overflow problems stop entirely — bf16 has the same range as FP32.
Follow-up questions to expect
- "Why not train entirely in bf16 with no FP32 master copy?" — Small updates round away, as in the
1.0 + 0.0001example. Pure bf16 training needs extra tricks such as stochastic rounding or compensated summation. - "Do you need loss scaling with bf16?" — No. bf16 has FP32's exponent range, so gradients do not underflow the way they do in FP16.
- "You get NaNs with bf16. Is precision the cause?" — Rarely. Look at the learning rate, bad data or exploding gradients first.