Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

What is gradient accumulation, and how does it enable training with small GPU memory?


Sixteen small passes, one optimizer stepMicro-batchof 2: forwardand backwardAdd gradientsto the bufferRepeat 16 timesOne optimizerstep,effective batch 32Zero the buffer20,000 examples divided by 32 is 625 optimizer steps per epoch.
Memory follows the micro-batch while the optimizer sees the big batch — you pay for it in time, not in GPU memory.

What you need to know

Why it saves memory

Training memory has two parts. Weights, gradients and optimizer states do not depend on batch size. Activations — the intermediate values saved for the backward pass — grow with micro-batch size times sequence length. Accumulation keeps the micro-batch small, so activations stay small, while the optimizer still sees a large, stable batch.

The config

Python
from trl import SFTConfigargs = SFTConfig(    output_dir="report-summariser",    per_device_train_batch_size=2,    gradient_accumulation_steps=16,     # effective batch 32 on one GPU    learning_rate=2e-4,                 # LoRA range; full fine-tuning uses ~1e-5 to 2e-5    lr_scheduler_type="cosine",    warmup_steps=0.03,                  # in Transformers 5, a value below 1 means a ratio    num_train_epochs=3,    bf16=True,    gradient_checkpointing=True,        # the TRL default; trades compute for memory)

Counting steps correctly

With 20,000 examples and an effective batch of 32, one epoch is 20,000 / 32 = 625 optimizer steps, and three epochs are 1,875. The learning-rate schedule, warmup, logging, evaluation and saving all count optimizer steps, not micro-batches. Warmup of 3% here is about 56 steps.

On 4 GPUs, the same effective batch of 32 is 2 x 4 x 4: micro-batch 2, accumulation 4. Keeping the effective batch fixed when the GPU count changes keeps results comparable.

What it does not do

  • It is not faster. Sixteen micro-batches cost about as much compute as one batch of 32 would, plus some loss of efficiency from small matrix shapes.
  • It is not exactly identical to a real big batch in every detail. With variable-length sequences, the loss must be normalised by the total number of tokens across the whole accumulation window. Transformers fixed a bug here in late 2024 (version 4.46); older code weighted short sequences too heavily. Transformers use LayerNorm, so the BatchNorm problem (statistics from small batches) does not apply.

Other memory tools that pair with it

Gradient checkpointing (recompute activations in the backward pass), a shorter max_length, packing or padding-free batches, 8-bit optimizers, and QLoRA.

A real-life example

A diagnostics company fine-tunes a medical-report summariser on a single 24 GB GPU. Reports are long — up to 3,000 tokens with the summary. A micro-batch of 8 runs out of memory; a micro-batch of 1 fits.

Their first attempt used micro-batch 1 with no accumulation. Each step's gradient came from one report, so the loss jumped around, and at learning rate 2e-4 the run became unstable. They switched to micro-batch 1 × 16 accumulation steps. Memory stayed the same, the loss curve became smooth, and the original learning rate worked. The epoch took about 10% longer than a hypothetical batch of 16 would have on a larger GPU — but that larger GPU was not available.

Follow-up questions to expect

  • "If I change accumulation, should I change the learning rate?" — If the effective batch changes, yes, retune it. If you only rebalance micro-batch against accumulation and keep the effective batch the same, no.
  • "Why not just use batch size 1?" — Gradients from one example are very noisy, which forces a lower learning rate and makes training less stable.
  • "How does it interact with multi-GPU training?" — Each GPU accumulates locally, and gradients are synchronised only at the optimizer step, which also saves communication.