Local LLM Deployment and Quantization

Fine-tuning Models Locally


You have an RTX 4090 with 24 GB of VRAM and 800 examples of how your support team writes replies. You load Llama 2 7B, call Trainer.train(), and the process dies before the first optimiser step completes. Not slowly, not with a warning — CUDA out of memory. Tried to allocate 26.94 GiB.

Nothing was misconfigured. Full fine-tuning a 7-billion-parameter model in mixed precision with Adam requires, per parameter: 2 bytes of FP16 weights, 2 bytes of FP16 gradients, 4 bytes of FP32 master weights, and 4 bytes each for Adam's first and second moment estimates. That is 16 bytes per parameter:

6.74×1096.74 \times 10^{9} × 16 = 1.08×10111.08 \times 10^{11} bytes = 108 GB, before a single activation is stored.

Your card is short by a factor of four and a half. Two 80 GB A100s would just about do it. And yet the same 800 examples, on the same 4090, train perfectly well in about forty minutes and produce a model that sounds like your support team. The technique that closes a 108 GB gap down to under 6 GB is worth understanding properly, because the same arithmetic decides every choice you make afterwards.

Why 7B will not train on 24 GB, and QLoRA willFull fine-tune, Llama 2 7B• Weights and gradients in fp16 — 27 GB• FP32 master weights — 27 GB more• Adam moments — another 54 GB• Asks for 108 GB, dies at 24 GBQLoRA on the same 4090• Base frozen in NF4 — 3.6 GB• Adapters are 0.59 pct of weights• Adam state for them — 320 MB• About 5.6 GB in all, headroom left
The out-of-memory error is not the model — it is the gradients, master copy and optimiser state, seven times the size of the fp16 weights.

Why local, and what it costs

Hosted fine-tuning is genuinely convenient, and for many teams it is the right call. But three of its costs are invisible until you hit them.

Your data leaves. Uploading support tickets, clinical notes, or contracts to a third party is a legal question, not a technical one. In regulated settings it is frequently the whole reason local training is on the table.

You rent the result. A hosted fine-tune usually lives on the provider's infrastructure. You pay per token to use it, often at a premium over the base model, and you cannot inspect it, quantise it, or run it offline. When the provider deprecates the base model — and they do, on their schedule — your fine-tune goes with it.

Iteration is slow and metered. Fine-tuning is an experimental process. You will try five data mixes and three hyperparameter sets before something works. On rented infrastructure each attempt has a price and a queue; on your own card the marginal cost of the sixth attempt is electricity.

Against that, local training costs you the setup, the debugging, and a real ceiling on model size. Nobody is full-fine-tuning a 70B on a desktop.

Four ways to change a model's behaviour

Before reaching for training at all, be clear about which tool the problem needs.

ApproachTrainable parameters (7B)Memory to trainGood atBad at
Prompting0NoneBehaviour you can describe in words; fast iterationConsistency over long runs; eats context on every call
Retrieval0NoneFacts that change; citing sourcesStyle, tone, output format
LoRA~40M (0.6%)~16 GBStyle, format, domain vocabulary, task specialisationTeaching genuinely new knowledge
QLoRA~40M (0.6%)~6 GBEverything LoRA does, on consumer hardwareSame limits, plus slightly slower per step
Full fine-tuning6.74B (100%)~108 GBDeep behavioural change, new languagesCost; catastrophic forgetting; needs far more data

Fine-tuning teaches a model how to say things. Retrieval teaches it what to say. Choosing the wrong one produces a model that confidently formats wrong answers beautifully.

LoRA, with the arithmetic

Low-Rank Adaptation starts from an observation about what fine-tuning actually does. If a weight matrix W∈Rd×kW \in \mathbb{R}^{d \times k} changes by ΔW\Delta W during training, that update turns out to be approximately low-rank — its information sits in a handful of directions, not in all dkdk entries.

So do not learn ΔW\Delta W directly. Factor it:

W′=W+αrBA,B∈Rd×r,  A∈Rr×k,  r≪min⁡(d,k)W' = W + \frac{\alpha}{r} BA, \qquad B \in \mathbb{R}^{d \times r},\; A \in \mathbb{R}^{r \times k}, \; r \ll \min(d,k)

WW stays frozen. Only AA and BB receive gradients. AA is initialised from a small random distribution and BB to zeros, so at step zero BA=0BA = 0 and the model is exactly the base model — training starts from a known-good point rather than a perturbed one.

Count the savings on Llama 2 7B — an older model, used for the arithmetic in this lesson because every attention projection in it is square, which keeps the counting simple: hidden size 4096, feed-forward intermediate 11008, 32 layers, rank r=16r = 16. Adapting every linear projection:

ProjectionShapeFull paramsLoRA params, r(din+dout)r(d_{in}+d_{out})
q, k, v, o (four)4096 × 409667.1M4 × 16 × 8192 = 524,288
gate, up (two)4096 × 1100890.2M2 × 16 × 15104 = 483,328
down (one)11008 × 409645.1M16 × 15104 = 241,664
Per layer202.4M1,249,280
32 layers6.48B39,976,960

Forty million trainable parameters, 0.59% of the model. Saved at FP16 that is an 80 MB adapter file — small enough to email, version in git-lfs, or keep twenty of on a laptop. Restrict LoRA to the query and value projections only, as the original paper did, and it falls to 8.4M parameters and a 17 MB file.

The optimiser state shrinks with it. Adam needs 8 bytes per trainable parameter, so 40M × 8 = 320 MB instead of 54 GB.

QLoRA: the frozen base gets quantised too

LoRA still keeps the frozen base weights in FP16 — 13.5 GB for a 7B — which is why the table above says about 16 GB. QLoRA removes that too, by storing the frozen base in a 4-bit datatype and dequantising each block on the fly during the forward and backward passes.

Three ideas make it work:

  • NF4, a 4-bit datatype whose sixteen levels are placed at the quantiles of a normal distribution rather than uniformly. Neural network weights are approximately Gaussian, so quantile spacing puts more resolution where the weights actually are.
  • Double quantisation: the per-block scale factors are themselves quantised, saving roughly a further 0.4 bits per parameter.
  • Paged optimisers, which spill optimiser state to CPU memory during gradient-checkpointing spikes instead of throwing an out-of-memory error.

Gradients flow through the frozen quantised weights to reach the adapters. The base is never updated, so its quantisation error is a fixed distortion the adapters learn around.

ComponentLoRA (FP16 base)QLoRA (NF4 base)
Frozen base weights13.5 GB3.6 GB
Adapter weights (40M, FP16)80 MB80 MB
Adapter gradients80 MB80 MB
Adam state (40M × 8 bytes)320 MB320 MB
Activations, checkpointed, seq 1024, batch 1~1.5 GB~1.5 GB
Total~15.5 GB~5.6 GB
Fits on24 GB card8 GB card

The cost is speed. Dequantising blocks on every forward pass makes QLoRA roughly 25–40% slower per step than LoRA. On a task that takes forty minutes, that is a fine trade for making it possible at all.

Preparing the data

Data quality determines the outcome far more than any hyperparameter. Most disappointing fine-tunes are dataset failures wearing a training-configuration disguise.

Format

Use JSONL with one conversation per line, in the messages format that every modern trainer accepts:

JSON
{"messages": [{"role": "system", "content": "You are a support agent for Acme Cloud. Be concise and never promise refunds."}, {"role": "user", "content": "My deploy has been queued for 40 minutes."}, {"role": "assistant", "content": "Queued deploys past 15 minutes almost always mean a build-cache lock. Run `acme deploy --purge-cache` and retry. If it queues again, send me the deploy ID and I will escalate."}]}

The system prompt must match what you will send at inference time. If you train with one system prompt and serve with another, you have trained the model to condition on a signal you then remove, and the behaviour will be inconsistent.

What good data looks like

RuleWhy, and what breaks if you ignore it
Consistent output format across every exampleThe model learns the distribution you show it. Mixed formats teach it to sample randomly between them
Cover the edge cases you care aboutA model trained only on clean questions will handle a malformed one by inventing something
Include refusals and "I don't know" responsesOtherwise every example teaches it that a confident answer always exists
Deduplicate aggressivelyNear-duplicates are the fastest route to memorisation instead of generalisation
Hold out 10–15% before you startWithout an untouched validation split you cannot tell learning from memorising
Mix in 10–20% general instruction dataNarrow training causes catastrophic forgetting: the model gets excellent at your task and forgets how to summarise a paragraph

On volume, the honest guidance is smaller than people expect:

GoalExamples needed
Tone and formatting only50–200
One well-defined task (classify, extract, reply)500–2,000
Domain vocabulary and reasoning patterns2,000–10,000
Genuinely new capability10,000+, and consider whether retrieval is the better tool

Two hundred carefully written examples beat twenty thousand scraped ones. The model imitates what you show it, including the mistakes.

Training the adapter

Bash
pip install -U torch transformers peft trl bitsandbytes datasets accelerate
Python
import torchfrom transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfigfrom peft import LoraConfigfrom trl import SFTTrainer, SFTConfigfrom datasets import load_datasetBASE = "meta-llama/Llama-3.1-8B-Instruct"   # any instruct model with a chat templatebnb = BitsAndBytesConfig(    load_in_4bit=True,    bnb_4bit_quant_type="nf4",           # quantile-spaced levels, not uniform    bnb_4bit_use_double_quant=True,      # quantise the scales as well    bnb_4bit_compute_dtype=torch.bfloat16,)model = AutoModelForCausalLM.from_pretrained(    BASE, quantization_config=bnb, device_map="auto")tok = AutoTokenizer.from_pretrained(BASE)peft_config = LoraConfig(    r=16,                                 # rank: 8 for style, 16 default, 32-64 for hard tasks    lora_alpha=32,                        # scaling is alpha/r = 2.0    lora_dropout=0.05,    target_modules=["q_proj", "k_proj", "v_proj", "o_proj",                    "gate_proj", "up_proj", "down_proj"],    task_type="CAUSAL_LM",)data = load_dataset("json", data_files={"train": "train.jsonl", "test": "val.jsonl"})trainer = SFTTrainer(    model=model,    train_dataset=data["train"],    eval_dataset=data["test"],    processing_class=tok,                 # applies the model's own chat template    peft_config=peft_config,    args=SFTConfig(        output_dir="./adapter",        num_train_epochs=3,        per_device_train_batch_size=1,        gradient_accumulation_steps=8,    # effective batch size 8        gradient_checkpointing=True,      # trades ~30% speed for a large memory saving        learning_rate=2e-4,               # 10-100x higher than full fine-tuning        lr_scheduler_type="cosine",        warmup_steps=0.03,                # a float below 1 = fraction of all steps        bf16=True,        max_length=1024,                  # was max_seq_length in older TRL        eval_strategy="epoch",        logging_steps=10,    ),)trainer.train()trainer.model.save_pretrained("./adapter")     # ~80-90 MB

The code trains Llama 3.1 8B Instruct — the model the rest of this course runs — rather than the Llama 2 7B of the arithmetic, because a support-reply dataset in messages format needs an instruct model that ships a chat template. Its numbers run somewhat higher than the tables: its 128K-token vocabulary keeps two embedding matrices of about 1 GB each in 16-bit, so budget roughly 8–9 GB of VRAM for QLoRA rather than 5.6 GB, which is still comfortable on a 12 GB card.

Three settings deserve explanation rather than copying. The learning rate of 2e-4 is far higher than full fine-tuning's 1e-5 or 2e-5, because you are training a small randomly-initialised module rather than nudging pretrained weights. gradient_accumulation_steps gives you a larger effective batch without the memory of a larger real batch — eight steps of batch 1 approximate one step of batch 8. And max_length multiplies activation memory directly; halving it from 2048 to 1024 is the quickest way out of an out-of-memory error.

Watch validation loss, not training loss. Training loss falling while validation loss rises is overfitting, and on datasets under a thousand examples it usually starts in epoch two or three.

From adapter to something you can run

The adapter is not a model. Getting it into a local runtime takes four steps, and each has a trap.

Merge into the full-precision base

Python
import torchfrom transformers import AutoModelForCausalLM, AutoTokenizerfrom peft import PeftModel# Critical: load the base in FP16, NOT 4-bit, even though you trained against 4-bitBASE = "meta-llama/Llama-3.1-8B-Instruct"base = AutoModelForCausalLM.from_pretrained(    BASE, dtype=torch.float16, device_map="cpu")merged = PeftModel.from_pretrained(base, "./adapter").merge_and_unload()merged.save_pretrained("./merged-fp16", safe_serialization=True)AutoTokenizer.from_pretrained(BASE).save_pretrained("./merged-fp16")

Merging computes W′=W+αrBAW' = W + \frac{\alpha}{r} BA once per adapted matrix and writes the result. With α=32\alpha = 32 and r=16r = 16 the scaling factor is 2.0.

The trap is merging into the 4-bit base. It runs without error and quietly bakes NF4 quantisation error into the weights, which you then quantise again to GGUF. Two rounds of lossy compression on the same tensors produce a model measurably worse than either alone, and the symptom — slightly degraded quality with no obvious cause — is miserable to diagnose. Always merge into FP16.

Test before converting

Python
from transformers import pipelinepipe = pipeline("text-generation", model="./merged-fp16", device_map="auto",                dtype=torch.float16)for q in ["My deploy has been queued for 40 minutes.",          "Can I get a refund for last month?",          # should refuse          "What is the capital of Portugal?"]:            # should still work    print(q, "->", pipe(q, max_new_tokens=120, do_sample=False)[0]["generated_text"])

Include an off-task question. If the model has forgotten how to answer general questions, that is catastrophic forgetting, and the fix is more general data in the mix or fewer epochs — not anything downstream.

Convert and quantise

Bash
cd llama.cpppython convert_hf_to_gguf.py ../merged-fp16 --outfile support-8b-f16.gguf --outtype f16./build/bin/llama-quantize support-8b-f16.gguf support-8b-Q4_K_M.gguf Q4_K_M./build/bin/llama-cli -m support-8b-Q4_K_M.gguf -p "My deploy is queued." -st -n 100

Budget disk space: 16 GB of merged FP16 safetensors, 16 GB of F16 GGUF, and 5 GB of output. Keep the F16 GGUF — requantising to a different level later starts from it, and regenerating it means redoing the merge.

Package for Ollama

Bash
cat > Modelfile <<'EOF'FROM ./support-8b-Q4_K_M.ggufPARAMETER temperature 0.3PARAMETER num_ctx 4096PARAMETER stop "<|eot_id|>"SYSTEM "You are a support agent for Acme Cloud. Be concise and never promise refunds."EOFollama create acme-support -f Modelfileollama run acme-support "My deploy has been queued for 40 minutes."

Put the same system prompt here that you trained with, and a stop token that matches the model family (<|eot_id|> for Llama 3). This is where the train/serve mismatch mentioned earlier actually happens to people.

The alternative: do not merge at all

llama.cpp can load a LoRA adapter alongside a base model, and Ollama's Modelfile has an ADAPTER directive. Converting the adapter to GGUF and keeping it separate means one 5 GB base file plus several 80–90 MB adapters, instead of one 5 GB file per fine-tune:

Bash
python convert_lora_to_gguf.py ./adapter --outfile acme-lora-f16.gguf./build/bin/llama-cli -m llama-3.1-8b-instruct-Q4_K_M.gguf --lora acme-lora-f16.gguf -p "..." -st

Merged models are marginally faster and simpler to distribute. Separate adapters are dramatically cheaper in disk when you maintain several fine-tunes of one base, and let you swap behaviour without reloading gigabytes.

Where fine-tunes go wrong

Using it to add facts. Fine-tuning on a product catalogue does not reliably install the catalogue. The model learns to produce text shaped like catalogue entries, then hallucinates plausible SKUs. Facts that change belong in a retrieval system; fine-tuning is for how the model behaves.

Wrong chat template. Every model family has its own special-token structure. Train Llama 3 with a Mistral template and the loss will fall convincingly while you teach the model to emit tokens the runtime does not recognise. Always use the tokeniser's own apply_chat_template.

Too many epochs on too little data. Three epochs on 500 examples is reasonable. Ten epochs on 500 examples produces a model that reproduces training examples verbatim and handles nothing else. Watch the validation curve.

Rank inflation. Rank 128 does not fix a bad dataset. It gives 320M trainable parameters, more overfitting, and a 640 MB adapter. Start at 16 and only raise it if validation loss plateaus above where you need it.

Skipping the merged-model test. When quality is disappointing after quantisation, you need to know whether the fine-tune, the merge, or the quantisation caused it. Testing at FP16 before conversion turns a three-way mystery into a one-line answer.

Fitting this into how you actually work

Treat the whole thing as a pipeline you can rerun, not a sequence of manual steps. Put data preparation, training, merging, conversion, and quantisation into one script with the dataset path as an argument. You will run it more times than you expect, and the runs that fail will fail somewhere in the middle.

Version the dataset alongside the adapter. An 80 MB adapter tells you nothing about how it was made; the JSONL that produced it tells you everything. When a fine-tune from three months ago starts behaving oddly against a newer base model, the dataset is what lets you retrain rather than guess.

Before training anything, spend an afternoon on the prompt. Write out fifty test cases and see how far a good system prompt with three worked examples gets you. Frequently it gets you to 90% of what the fine-tune would have achieved, in an afternoon rather than a week, and with the enormous advantage that you can change it by editing a string. Fine-tune when prompting has plateaued, when the prompt has grown so long it costs real latency on every call, or when you need consistency that instructions cannot enforce — those are the cases where forty minutes on a consumer GPU genuinely pays for itself.

Finally, size the job before you start it. Multiply parameters by 16 bytes for a full fine-tune, or by 0.5 bytes plus roughly 500 MB of adapter overhead for QLoRA, and compare against your card. That two-minute calculation is the difference between an evening of training and an evening of reading out-of-memory tracebacks.