Course Content
LLMOps & Deployment
6 sections · 40 lessons
What is LoRA, and why is it useful for efficient fine-tuning?
What you need to know
The numbers
W_effective = W_frozen + (B · A) × (alpha / r)One 4096 × 4096 matrix: 16,777,216 weightsLoRA with r = 16: 2 × 4096 × 16 = 131,072 (0.78%)Applying rank-16 LoRA to all attention and MLP projections of Llama 3.1 8B trains about 42 million parameters — 0.5% of the model — and the adapter is about 84 MB in bf16, against 16 GB for the full model. alpha is a scaling factor that controls how strongly the update is applied.
Why it saves so much memory
Full fine-tuning keeps gradients and optimizer state (Adam keeps two extra numbers per weight) for all 8 billion parameters — well over 100 GB in total. LoRA keeps them only for the adapter. QLoRA also stores the frozen base in 4-bit (NF4), so a 70B model can be fine-tuned on a single 80 GB GPU.
Why operations teams like it
Many adapters, one base. vLLM loads the base once and applies adapters per request:
1vllm serve meta-llama/Llama-3.1-8B-Instruct \2 --enable-lora \3 --lora-modules pensions=/models/lora/pensions housing=/models/lora/housing \4 --max-loras 4 --max-lora-rank 16A request with "model": "pensions" uses that adapter, so ten fine-tunes do not need ten GPUs. --max-loras is how many adapters can be active in one batch.
- Cheap rollback — point back to the previous adapter.
- Merge when there is only one — adding
B·AintoWremoves all runtime overhead.
What LoRA is good and bad at
Good: output format, tone, domain vocabulary, following a house style, a narrow task. Weak: teaching large amounts of new facts — the model may learn to sound like it knows them. Facts that change belong in retrieval.
A real-life example
A state government's chatbot serves eight departments. Each wants its own tone and answer format: pensions wants step-by-step answers with form numbers; housing wants a short eligibility check first; health must always add a helpline number.
Instead of eight fine-tuned models (8 × 16 GB, eight GPUs), the team trains eight rank-16 LoRA adapters on one 8B base, each on 2,000–4,000 reviewed conversations. Each adapter trains in under two hours on one GPU and is about 84 MB. One vLLM replica per GPU serves the base plus all adapters; the router picks the adapter by department.
When the housing adapter starts adding an outdated income limit (a fact it learned from training data), the team rolls back to the previous adapter in seconds and moves the income limits into the retrieval index, where they can be updated the day a rule changes.
Follow-up questions to expect
- "How do you choose the rank?" — Start at 8–16; raise it if the task is complex and the eval score keeps rising. Very high ranks approach full fine-tuning cost.
- "Which layers get adapters?" — Commonly all attention and MLP projections; attention-only is cheaper but often a little weaker.
- "Does LoRA slow inference?" — Slightly when unmerged and serving many adapters; zero when merged into the base.