Course Content
Fine-Tuning LLMs
6 sections · 52 lessons
For a domain assistant, how do you choose between LoRA and full fine-tuning?
What you need to know
The decision in one table
| Question | Points to LoRA | Points to full fine-tuning |
|---|---|---|
| Does the base understand the domain's words? | Yes | No — new language, notation or modality |
| How much data? | Hundreds to tens of thousands of examples | Hundreds of millions of tokens |
| How many variants? | Several teams or customers | Exactly one |
| Hardware | One GPU | A multi-GPU node with sharding |
| Iteration speed | Weekly changes | Rare, planned releases |
The cost gap for an 8B model
Text
Full fine-tuning: ~128 GB for weights, gradients, Adam -> 2-4 x 80 GB GPUs + shardingLoRA (bf16 base): ~17 GB before activations -> one 24-48 GB GPUServing: full FT = a separate 16 GB model per variant LoRA = one base + an ~84 MB adapter per variantTune LoRA properly before giving up on it
- Target all linear layers, including the MLP.
- Try ranks 16, 64 and 128 — with
use_rslora=Trueat high ranks. - Sweep the learning rate; LoRA's best rate is usually about 10 times the full fine-tuning rate.
- Train long enough, and check the data first.
Recent studies found LoRA matches full fine-tuning on small-to-medium post-training datasets when set up this way. The gap appears with very large datasets and big distribution shifts — exactly the "full fine-tuning" column above.
A real-life example
A diagnostics chain builds a radiology-report assistant on an 8B model. Its steps:
- LoRA r=16, attention only: 71% of summaries accepted by radiologists without edits.
- LoRA r=64, all linear layers, learning-rate sweep: 78%.
- Full fine-tuning on a rented 8-GPU node: 78.5%, inside the noise of a 400-report test set.
Error analysis shows most remaining failures come from 300 training summaries written by a vendor in an old template. Fixing those gives 83% with the LoRA setup. The "we need full fine-tuning" conclusion had really been "our data has a bad source".
Follow-up questions to expect
- "When is full fine-tuning clearly right?" — Continued pretraining on billions of tokens, a new language or modality, or a foundation-model team building its own instruct model.
- "Does full fine-tuning forget more?" — Usually yes, because every weight can move. Plan for replay data and regression tests.
- "What about QLoRA?" — Same decision as LoRA when memory is tight; it trades speed for memory, not quality in most reports.