Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

What is LoRA (Low-Rank Adaptation) and what changes during training?


One 4096 by 4096 projection with a rank-16 LoRAInput x:4,096 numbersA squeezesto 16 numbersB expandsback to 4,096Scale byalpha over rAdd tofrozen W xA and B hold 131,072 weights; W holds 16.8 million and never changes.
B starts at zero, so step 0 is exactly the base model, and everything learned lives in 0.78% of the matrix's size.

What you need to know

The formula in words

Text
original:  h = W xwith LoRA: h = W x + (alpha / r) * B (A x)W: d_out x d_in   (frozen)A: r x d_in       (trained, starts small and random)B: d_out x r      (trained, starts at zero)

A squeezes the input down to r numbers, and B expands them back to the output size. Because B starts at zero, B A x is zero at step 0, so the model starts exactly as the base model. A is random, so the gradient reaching B is not zero and learning can start. (If both started at zero, neither would ever receive a gradient.)

Worked arithmetic on a 4096 × 4096 matrix

LoRA parameters = r × (d_in + d_out) = r × 8,192. The full matrix has 16,777,216.

Rank rLoRA paramsShare of the matrix
432,7680.20%
865,5360.39%
16131,0720.78%
64524,2883.1%

Across a whole model the numbers add up. On Llama-3.1-8B at r = 16, targeting only the four attention projections gives 13.6M trainable parameters (0.17%); targeting all seven linear layers per block, including the MLP, gives 41.9M (0.52%).

Code (peft 0.21, transformers 5.x)

Python
import torchfrom transformers import AutoModelForCausalLMfrom peft import LoraConfig, get_peft_modelbase = AutoModelForCausalLM.from_pretrained(    "meta-llama/Llama-3.1-8B-Instruct", dtype=torch.bfloat16)config = LoraConfig(    r=16, lora_alpha=32, lora_dropout=0.05,    target_modules="all-linear",   # q,k,v,o and gate,up,down    task_type="CAUSAL_LM")model = get_peft_model(base, config)model.print_trainable_parameters()# trainable params: 41,943,040 || all params: 8,072,204,288 || trainable%: 0.5196

r sets capacity. lora_alpha sets the strength: the update is multiplied by alpha / r, here 2. Common defaults are alpha = r or alpha = 2r. use_rslora=True scales by alpha / √r instead, which keeps training stable when you use large ranks. use_dora=True switches to DoRA, which also learns a magnitude per output channel (1.4M extra parameters on this model).

After training

model.merge_and_unload() adds (alpha / r) * B A into W and removes the adapter, so inference is as fast as the base model. Or keep adapters separate and serve many of them over one base (section 4 covers serving).

A real-life example

A Mumbai law firm reviews hundreds of commercial contracts a month. It wants each clause labelled — indemnity, limitation of liability, arbitration seat, governing law, termination, confidentiality — with a one-line risk note.

The team trains a rank-16, all-linear LoRA adapter on 6,000 clauses labelled by associates. The adapter file is about 84 MB, while the base model is 16 GB. Later, the firm wants a second behaviour — drafting fallback clauses — so it trains a second adapter over the same base instead of a second model. The frozen base means the model's general English and reasoning stay intact, and the firm can prove which adapter version produced any label.

Follow-up questions to expect

  • "Which modules should you target?" — All linear layers usually beat attention-only at the same budget, because much of the knowledge lives in the MLP layers. target_modules="all-linear" does this in PEFT.
  • "What does lora_alpha do?" — It scales the update. If you double r but keep alpha fixed, each direction's contribution halves, so re-tune alpha or the learning rate when you change rank.
  • "Does LoRA add inference latency?" — Only if unmerged: an extra small matrix multiply per targeted layer. Merged, it adds nothing.
  • "What learning rate?" — LoRA typically uses a learning rate around 10 times higher than full fine-tuning, for example 1e-4 to 2e-4 instead of 1e-5 to 2e-5.