Course Content
Fine-Tuning LLMs
6 sections · 52 lessons
In production, how can you host and switch between multiple LoRA adapters?
What you need to know
Why this works
A LoRA layer computes y = W·x + B·(A·x). W is the big frozen base matrix and is the same for every request. A and B are the small adapter matrices, and only they change between tenants. So the server runs the expensive W·x once for the whole batch and adds a cheap, different low-rank term for each row.
The adapters are tiny. A rank-16 LoRA on all linear layers of Llama-3.1-8B has about 42 million parameters, which is about 84 MB in bf16. The base model is about 16 GB. Special batched kernels (from the 2023 Punica and S-LoRA work, now built into vLLM) apply a different adapter to each row of a batch in one GPU call, so mixing tenants in one batch does not make the server run them one by one.
Serving with vLLM
1# vLLM 0.30 (the current release at the time of writing)2vllm serve meta-llama/Llama-3.1-8B-Instruct \3 --enable-lora \4 --max-loras 8 --max-lora-rank 16 --max-cpu-loras 40 \5 --lora-modules brand-a-v7=/adapters/brand-a/v7 brand-b-v3=/adapters/brand-b/v31from openai import OpenAI2client = OpenAI(base_url="http://localhost:8000/v1", api_key="unused")3reply = client.chat.completions.create(4 model="brand-a-v7", # the adapter name picks the fine-tune5 messages=[{"role": "user", "content": "Mera order kab aayega?"}],6)The client code is the normal OpenAI-compatible API. The only change is that model names an adapter instead of the base model.
| Setting | What it controls |
|---|---|
--max-loras | How many different adapters can be in one running batch (GPU slots). |
--max-lora-rank | The largest rank allowed. Memory is reserved for this rank, so set it to your real maximum, not "just in case" 256. |
--max-cpu-loras | How many adapters are cached in CPU memory. Adapters move between CPU and GPU slots in least-recently-used order. |
| Runtime loading | With VLLM_ALLOW_RUNTIME_LORA_UPDATING=True, adapters can be added through /v1/load_lora_adapter. vLLM's docs warn this is not safe to expose; keep it admin-only. |
SGLang and LoRAX also serve many adapters on one base, and several hosted platforms offer this as a managed service.
Operating rules
- Same base, always. An adapter is a correction to specific weights. Store the base model's name and revision (or file hash) in the adapter metadata and refuse to load a mismatch.
- Immutable versions. Name adapters
brand-a-v7, neverbrand-a-latest. The router maps a tenant to a version, so rollout and rollback are config changes. - Measure the overhead. Unmerged adapters cost a little extra per token, more with higher rank and more distinct adapters per batch. Measure it at your real traffic mix.
- Merge the hot path. If one adapter carries most of the traffic and latency matters, merge it into the base (
merge_and_unload()in PEFT) and give it its own deployment. Keep the multi-LoRA server for the long tail.
A real-life example
A customer-support platform in Pune serves 35 D2C brands. Each brand has its own Hindi-and-English support adapter (rank 16, about 84 MB) trained on the same Llama-3.1-8B-Instruct base.
Separate deployments would need 35 copies of a 16 GB model — 560 GB of weights, or seven H100s before any KV cache. Instead, one H100 (80 GB) holds the base once (16 GB), up to 8 adapters on the GPU, 35 adapters (about 3 GB) cached in CPU memory, and uses the rest for KV cache. A small router maps tenant_id to an adapter name.
When brand A's new adapter brand-a-v8 is ready, the team registers it next to v7 and sends 10% of brand A's traffic to it. After a week, escalations to human agents are flat and customer ratings are slightly up, so the router switches fully. One brand carries 40% of all traffic; its adapter is merged and served on a dedicated GPU, which removes the adapter overhead on the busiest path.
Follow-up questions to expect
- "What happens when more adapters are active than
--max-loras?" — Requests for an adapter without a GPU slot wait until one frees up; the least recently used adapter is swapped out. Set--max-lorasto the number of distinct adapters you expect in one batch. - "Can adapters have different ranks?" — Yes, up to
--max-lora-rank. But memory is reserved for the maximum, so one rank-128 adapter raises the cost for everyone. - "Can a QLoRA-trained adapter be served on a bf16 base?" — Usually yes, and it is the normal path. The adapter was trained against slightly different (dequantized) weights, so run your eval on the served setup.
- "Why not merge every adapter?" — Each merged model is a full 16 GB copy. Merging is for a few hot tenants, not fifty.