Course Content
Local LLM Deployment and Quantization
3 sections · 7 lessons
Fine-tuning Models Locally
You have an RTX 4090 with 24 GB of VRAM and 800 examples of how your support team writes replies. You load Llama 2 7B, call Trainer.train(), and the process dies before the first optimiser step completes. Not slowly, not with a warning — CUDA out of memory. Tried to allocate 26.94 GiB.
Nothing was misconfigured. Full fine-tuning a 7-billion-parameter model in mixed precision with Adam requires, per parameter: 2 bytes of FP16 weights, 2 bytes of FP16 gradients, 4 bytes of FP32 master weights, and 4 bytes each for Adam's first and second moment estimates. That is 16 bytes per parameter:
6.74×109 × 16 = 1.08×1011 bytes = 108 GB, before a single activation is stored.
Your card is short by a factor of four and a half. Two 80 GB A100s would just about do it. And yet the same 800 examples, on the same 4090, train perfectly well in about forty minutes and produce a model that sounds like your support team. The technique that closes a 108 GB gap down to under 6 GB is worth understanding properly, because the same arithmetic decides every choice you make afterwards.
Why local, and what it costs
Hosted fine-tuning is genuinely convenient, and for many teams it is the right call. But three of its costs are invisible until you hit them.
Your data leaves. Uploading support tickets, clinical notes, or contracts to a third party is a legal question, not a technical one. In regulated settings it is frequently the whole reason local training is on the table.
You rent the result. A hosted fine-tune usually lives on the provider's infrastructure. You pay per token to use it, often at a premium over the base model, and you cannot inspect it, quantise it, or run it offline. When the provider deprecates the base model — and they do, on their schedule — your fine-tune goes with it.
Iteration is slow and metered. Fine-tuning is an experimental process. You will try five data mixes and three hyperparameter sets before something works. On rented infrastructure each attempt has a price and a queue; on your own card the marginal cost of the sixth attempt is electricity.
Against that, local training costs you the setup, the debugging, and a real ceiling on model size. Nobody is full-fine-tuning a 70B on a desktop.
Four ways to change a model's behaviour
Before reaching for training at all, be clear about which tool the problem needs.
| Approach | Trainable parameters (7B) | Memory to train | Good at | Bad at |
|---|---|---|---|---|
| Prompting | 0 | None | Behaviour you can describe in words; fast iteration | Consistency over long runs; eats context on every call |
| Retrieval | 0 | None | Facts that change; citing sources | Style, tone, output format |
| LoRA | ~40M (0.6%) | ~16 GB | Style, format, domain vocabulary, task specialisation | Teaching genuinely new knowledge |
| QLoRA | ~40M (0.6%) | ~6 GB | Everything LoRA does, on consumer hardware | Same limits, plus slightly slower per step |
| Full fine-tuning | 6.74B (100%) | ~108 GB | Deep behavioural change, new languages | Cost; catastrophic forgetting; needs far more data |
Fine-tuning teaches a model how to say things. Retrieval teaches it what to say. Choosing the wrong one produces a model that confidently formats wrong answers beautifully.
LoRA, with the arithmetic
Low-Rank Adaptation starts from an observation about what fine-tuning actually does. If a weight matrix W∈Rd×k changes by ΔW during training, that update turns out to be approximately low-rank — its information sits in a handful of directions, not in all dk entries.
So do not learn ΔW directly. Factor it:
W stays frozen. Only A and B receive gradients. A is initialised from a small random distribution and B to zeros, so at step zero BA=0 and the model is exactly the base model — training starts from a known-good point rather than a perturbed one.
Count the savings on Llama 2 7B — an older model, used for the arithmetic in this lesson because every attention projection in it is square, which keeps the counting simple: hidden size 4096, feed-forward intermediate 11008, 32 layers, rank r=16. Adapting every linear projection:
| Projection | Shape | Full params | LoRA params, r(din+dout) |
|---|---|---|---|
| q, k, v, o (four) | 4096 × 4096 | 67.1M | 4 × 16 × 8192 = 524,288 |
| gate, up (two) | 4096 × 11008 | 90.2M | 2 × 16 × 15104 = 483,328 |
| down (one) | 11008 × 4096 | 45.1M | 16 × 15104 = 241,664 |
| Per layer | 202.4M | 1,249,280 | |
| 32 layers | 6.48B | 39,976,960 |
Forty million trainable parameters, 0.59% of the model. Saved at FP16 that is an 80 MB adapter file — small enough to email, version in git-lfs, or keep twenty of on a laptop. Restrict LoRA to the query and value projections only, as the original paper did, and it falls to 8.4M parameters and a 17 MB file.
The optimiser state shrinks with it. Adam needs 8 bytes per trainable parameter, so 40M × 8 = 320 MB instead of 54 GB.
QLoRA: the frozen base gets quantised too
LoRA still keeps the frozen base weights in FP16 — 13.5 GB for a 7B — which is why the table above says about 16 GB. QLoRA removes that too, by storing the frozen base in a 4-bit datatype and dequantising each block on the fly during the forward and backward passes.
Three ideas make it work:
- NF4, a 4-bit datatype whose sixteen levels are placed at the quantiles of a normal distribution rather than uniformly. Neural network weights are approximately Gaussian, so quantile spacing puts more resolution where the weights actually are.
- Double quantisation: the per-block scale factors are themselves quantised, saving roughly a further 0.4 bits per parameter.
- Paged optimisers, which spill optimiser state to CPU memory during gradient-checkpointing spikes instead of throwing an out-of-memory error.
Gradients flow through the frozen quantised weights to reach the adapters. The base is never updated, so its quantisation error is a fixed distortion the adapters learn around.
| Component | LoRA (FP16 base) | QLoRA (NF4 base) |
|---|---|---|
| Frozen base weights | 13.5 GB | 3.6 GB |
| Adapter weights (40M, FP16) | 80 MB | 80 MB |
| Adapter gradients | 80 MB | 80 MB |
| Adam state (40M × 8 bytes) | 320 MB | 320 MB |
| Activations, checkpointed, seq 1024, batch 1 | ~1.5 GB | ~1.5 GB |
| Total | ~15.5 GB | ~5.6 GB |
| Fits on | 24 GB card | 8 GB card |
The cost is speed. Dequantising blocks on every forward pass makes QLoRA roughly 25–40% slower per step than LoRA. On a task that takes forty minutes, that is a fine trade for making it possible at all.
Preparing the data
Data quality determines the outcome far more than any hyperparameter. Most disappointing fine-tunes are dataset failures wearing a training-configuration disguise.
Format
Use JSONL with one conversation per line, in the messages format that every modern trainer accepts:
{"messages": [{"role": "system", "content": "You are a support agent for Acme Cloud. Be concise and never promise refunds."}, {"role": "user", "content": "My deploy has been queued for 40 minutes."}, {"role": "assistant", "content": "Queued deploys past 15 minutes almost always mean a build-cache lock. Run `acme deploy --purge-cache` and retry. If it queues again, send me the deploy ID and I will escalate."}]}The system prompt must match what you will send at inference time. If you train with one system prompt and serve with another, you have trained the model to condition on a signal you then remove, and the behaviour will be inconsistent.
What good data looks like
| Rule | Why, and what breaks if you ignore it |
|---|---|
| Consistent output format across every example | The model learns the distribution you show it. Mixed formats teach it to sample randomly between them |
| Cover the edge cases you care about | A model trained only on clean questions will handle a malformed one by inventing something |
| Include refusals and "I don't know" responses | Otherwise every example teaches it that a confident answer always exists |
| Deduplicate aggressively | Near-duplicates are the fastest route to memorisation instead of generalisation |
| Hold out 10–15% before you start | Without an untouched validation split you cannot tell learning from memorising |
| Mix in 10–20% general instruction data | Narrow training causes catastrophic forgetting: the model gets excellent at your task and forgets how to summarise a paragraph |
On volume, the honest guidance is smaller than people expect:
| Goal | Examples needed |
|---|---|
| Tone and formatting only | 50–200 |
| One well-defined task (classify, extract, reply) | 500–2,000 |
| Domain vocabulary and reasoning patterns | 2,000–10,000 |
| Genuinely new capability | 10,000+, and consider whether retrieval is the better tool |
Two hundred carefully written examples beat twenty thousand scraped ones. The model imitates what you show it, including the mistakes.
Training the adapter
pip install -U torch transformers peft trl bitsandbytes datasets accelerate1import torch2from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig3from peft import LoraConfig4from trl import SFTTrainer, SFTConfig5from datasets import load_dataset67BASE = "meta-llama/Llama-3.1-8B-Instruct" # any instruct model with a chat template89bnb = BitsAndBytesConfig(10 load_in_4bit=True,11 bnb_4bit_quant_type="nf4", # quantile-spaced levels, not uniform12 bnb_4bit_use_double_quant=True, # quantise the scales as well13 bnb_4bit_compute_dtype=torch.bfloat16,14)1516model = AutoModelForCausalLM.from_pretrained(17 BASE, quantization_config=bnb, device_map="auto")18tok = AutoTokenizer.from_pretrained(BASE)1920peft_config = LoraConfig(21 r=16, # rank: 8 for style, 16 default, 32-64 for hard tasks22 lora_alpha=32, # scaling is alpha/r = 2.023 lora_dropout=0.05,24 target_modules=["q_proj", "k_proj", "v_proj", "o_proj",25 "gate_proj", "up_proj", "down_proj"],26 task_type="CAUSAL_LM",27)2829data = load_dataset("json", data_files={"train": "train.jsonl", "test": "val.jsonl"})3031trainer = SFTTrainer(32 model=model,33 train_dataset=data["train"],34 eval_dataset=data["test"],35 processing_class=tok, # applies the model's own chat template36 peft_config=peft_config,37 args=SFTConfig(38 output_dir="./adapter",39 num_train_epochs=3,40 per_device_train_batch_size=1,41 gradient_accumulation_steps=8, # effective batch size 842 gradient_checkpointing=True, # trades ~30% speed for a large memory saving43 learning_rate=2e-4, # 10-100x higher than full fine-tuning44 lr_scheduler_type="cosine",45 warmup_steps=0.03, # a float below 1 = fraction of all steps46 bf16=True,47 max_length=1024, # was max_seq_length in older TRL48 eval_strategy="epoch",49 logging_steps=10,50 ),51)52trainer.train()53trainer.model.save_pretrained("./adapter") # ~80-90 MBThe code trains Llama 3.1 8B Instruct — the model the rest of this course runs — rather than the Llama 2 7B of the arithmetic, because a support-reply dataset in messages format needs an instruct model that ships a chat template. Its numbers run somewhat higher than the tables: its 128K-token vocabulary keeps two embedding matrices of about 1 GB each in 16-bit, so budget roughly 8–9 GB of VRAM for QLoRA rather than 5.6 GB, which is still comfortable on a 12 GB card.
Three settings deserve explanation rather than copying. The learning rate of 2e-4 is far higher than full fine-tuning's 1e-5 or 2e-5, because you are training a small randomly-initialised module rather than nudging pretrained weights. gradient_accumulation_steps gives you a larger effective batch without the memory of a larger real batch — eight steps of batch 1 approximate one step of batch 8. And max_length multiplies activation memory directly; halving it from 2048 to 1024 is the quickest way out of an out-of-memory error.
Watch validation loss, not training loss. Training loss falling while validation loss rises is overfitting, and on datasets under a thousand examples it usually starts in epoch two or three.
From adapter to something you can run
The adapter is not a model. Getting it into a local runtime takes four steps, and each has a trap.
Merge into the full-precision base
1import torch2from transformers import AutoModelForCausalLM, AutoTokenizer3from peft import PeftModel45# Critical: load the base in FP16, NOT 4-bit, even though you trained against 4-bit6BASE = "meta-llama/Llama-3.1-8B-Instruct"7base = AutoModelForCausalLM.from_pretrained(8 BASE, dtype=torch.float16, device_map="cpu")910merged = PeftModel.from_pretrained(base, "./adapter").merge_and_unload()11merged.save_pretrained("./merged-fp16", safe_serialization=True)12AutoTokenizer.from_pretrained(BASE).save_pretrained("./merged-fp16")Merging computes W′=W+rαBA once per adapted matrix and writes the result. With α=32 and r=16 the scaling factor is 2.0.
The trap is merging into the 4-bit base. It runs without error and quietly bakes NF4 quantisation error into the weights, which you then quantise again to GGUF. Two rounds of lossy compression on the same tensors produce a model measurably worse than either alone, and the symptom — slightly degraded quality with no obvious cause — is miserable to diagnose. Always merge into FP16.
Test before converting
1from transformers import pipeline2pipe = pipeline("text-generation", model="./merged-fp16", device_map="auto",3 dtype=torch.float16)45for q in ["My deploy has been queued for 40 minutes.",6 "Can I get a refund for last month?", # should refuse7 "What is the capital of Portugal?"]: # should still work8 print(q, "->", pipe(q, max_new_tokens=120, do_sample=False)[0]["generated_text"])Include an off-task question. If the model has forgotten how to answer general questions, that is catastrophic forgetting, and the fix is more general data in the mix or fewer epochs — not anything downstream.
Convert and quantise
1cd llama.cpp2python convert_hf_to_gguf.py ../merged-fp16 --outfile support-8b-f16.gguf --outtype f163./build/bin/llama-quantize support-8b-f16.gguf support-8b-Q4_K_M.gguf Q4_K_M4./build/bin/llama-cli -m support-8b-Q4_K_M.gguf -p "My deploy is queued." -st -n 100Budget disk space: 16 GB of merged FP16 safetensors, 16 GB of F16 GGUF, and 5 GB of output. Keep the F16 GGUF — requantising to a different level later starts from it, and regenerating it means redoing the merge.
Package for Ollama
1cat > Modelfile <<'EOF'2FROM ./support-8b-Q4_K_M.gguf3PARAMETER temperature 0.34PARAMETER num_ctx 40965PARAMETER stop "<|eot_id|>"6SYSTEM "You are a support agent for Acme Cloud. Be concise and never promise refunds."7EOF89ollama create acme-support -f Modelfile10ollama run acme-support "My deploy has been queued for 40 minutes."Put the same system prompt here that you trained with, and a stop token that matches the model family (<|eot_id|> for Llama 3). This is where the train/serve mismatch mentioned earlier actually happens to people.
The alternative: do not merge at all
llama.cpp can load a LoRA adapter alongside a base model, and Ollama's Modelfile has an ADAPTER directive. Converting the adapter to GGUF and keeping it separate means one 5 GB base file plus several 80–90 MB adapters, instead of one 5 GB file per fine-tune:
python convert_lora_to_gguf.py ./adapter --outfile acme-lora-f16.gguf./build/bin/llama-cli -m llama-3.1-8b-instruct-Q4_K_M.gguf --lora acme-lora-f16.gguf -p "..." -stMerged models are marginally faster and simpler to distribute. Separate adapters are dramatically cheaper in disk when you maintain several fine-tunes of one base, and let you swap behaviour without reloading gigabytes.
Where fine-tunes go wrong
Using it to add facts. Fine-tuning on a product catalogue does not reliably install the catalogue. The model learns to produce text shaped like catalogue entries, then hallucinates plausible SKUs. Facts that change belong in a retrieval system; fine-tuning is for how the model behaves.
Wrong chat template. Every model family has its own special-token structure. Train Llama 3 with a Mistral template and the loss will fall convincingly while you teach the model to emit tokens the runtime does not recognise. Always use the tokeniser's own apply_chat_template.
Too many epochs on too little data. Three epochs on 500 examples is reasonable. Ten epochs on 500 examples produces a model that reproduces training examples verbatim and handles nothing else. Watch the validation curve.
Rank inflation. Rank 128 does not fix a bad dataset. It gives 320M trainable parameters, more overfitting, and a 640 MB adapter. Start at 16 and only raise it if validation loss plateaus above where you need it.
Skipping the merged-model test. When quality is disappointing after quantisation, you need to know whether the fine-tune, the merge, or the quantisation caused it. Testing at FP16 before conversion turns a three-way mystery into a one-line answer.
Fitting this into how you actually work
Treat the whole thing as a pipeline you can rerun, not a sequence of manual steps. Put data preparation, training, merging, conversion, and quantisation into one script with the dataset path as an argument. You will run it more times than you expect, and the runs that fail will fail somewhere in the middle.
Version the dataset alongside the adapter. An 80 MB adapter tells you nothing about how it was made; the JSONL that produced it tells you everything. When a fine-tune from three months ago starts behaving oddly against a newer base model, the dataset is what lets you retrain rather than guess.
Before training anything, spend an afternoon on the prompt. Write out fifty test cases and see how far a good system prompt with three worked examples gets you. Frequently it gets you to 90% of what the fine-tune would have achieved, in an afternoon rather than a week, and with the enormous advantage that you can change it by editing a string. Fine-tune when prompting has plateaued, when the prompt has grown so long it costs real latency on every call, or when you need consistency that instructions cannot enforce — those are the cases where forty minutes on a consumer GPU genuinely pays for itself.
Finally, size the job before you start it. Multiply parameters by 16 bytes for a full fine-tune, or by 0.5 bytes plus roughly 500 MB of adapter overhead for QLoRA, and compare against your card. That two-minute calculation is the difference between an evening of training and an evening of reading out-of-memory tracebacks.