Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

For a domain assistant, how do you choose between LoRA and full fine-tuning?


Radiologist acceptance rate, one change at a time717878.5830123LoRA r16,attentionLoRA r64,all-linearFull fine-tuneLoRA + fixed dataFull fine-tuning's extra half point was inside the noise of a 400-report test set.
Tuning LoRA properly and fixing one bad data source beat buying an 8-GPU node.

What you need to know

The decision in one table

QuestionPoints to LoRAPoints to full fine-tuning
Does the base understand the domain's words?YesNo — new language, notation or modality
How much data?Hundreds to tens of thousands of examplesHundreds of millions of tokens
How many variants?Several teams or customersExactly one
HardwareOne GPUA multi-GPU node with sharding
Iteration speedWeekly changesRare, planned releases

The cost gap for an 8B model

Text
Full fine-tuning: ~128 GB for weights, gradients, Adam  -> 2-4 x 80 GB GPUs + shardingLoRA (bf16 base): ~17 GB before activations            -> one 24-48 GB GPUServing:          full FT = a separate 16 GB model per variant                  LoRA    = one base + an ~84 MB adapter per variant

Tune LoRA properly before giving up on it

  • Target all linear layers, including the MLP.
  • Try ranks 16, 64 and 128 — with use_rslora=True at high ranks.
  • Sweep the learning rate; LoRA's best rate is usually about 10 times the full fine-tuning rate.
  • Train long enough, and check the data first.

Recent studies found LoRA matches full fine-tuning on small-to-medium post-training datasets when set up this way. The gap appears with very large datasets and big distribution shifts — exactly the "full fine-tuning" column above.

A real-life example

A diagnostics chain builds a radiology-report assistant on an 8B model. Its steps:

  1. LoRA r=16, attention only: 71% of summaries accepted by radiologists without edits.
  2. LoRA r=64, all linear layers, learning-rate sweep: 78%.
  3. Full fine-tuning on a rented 8-GPU node: 78.5%, inside the noise of a 400-report test set.

Error analysis shows most remaining failures come from 300 training summaries written by a vendor in an old template. Fixing those gives 83% with the LoRA setup. The "we need full fine-tuning" conclusion had really been "our data has a bad source".

Follow-up questions to expect

  • "When is full fine-tuning clearly right?" — Continued pretraining on billions of tokens, a new language or modality, or a foundation-model team building its own instruct model.
  • "Does full fine-tuning forget more?" — Usually yes, because every weight can move. Plan for replay data and regression tests.
  • "What about QLoRA?" — Same decision as LoRA when memory is tight; it trades speed for memory, not quality in most reports.