Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

What does adapter-based fine-tuning mean in practice?


What you need to know

Bottleneck adapters vs LoRA adapters

A bottleneck adapter is a tiny network inserted after a layer: project down from 4,096 to, say, 64 dimensions, apply a non-linearity, project back up to 4,096, and add the result to the input. That is about 2 × 4,096 × 64 ≈ 524,000 parameters per adapter.

Bottleneck adapter

  • Extra layer in the forward path
  • Adds latency every token
  • Cannot be merged into the weights
  • Non-linear, so slightly more expressive

LoRA adapter

  • Runs beside an existing matrix
  • Can be merged: zero extra latency
  • Can be served unmerged, many per base
  • The default in 2026

The life of an adapter

Python
from transformers import AutoModelForCausalLMfrom peft import PeftModelbase = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-7B-Instruct")model = PeftModel.from_pretrained(base, "adapters/legal-v3", adapter_name="legal")model.load_adapter("adapters/brand-v2", adapter_name="brand")model.set_adapter("brand")            # switch behaviour, no reloadwith model.disable_adapter():         # plain base model, e.g. for comparison    ...merged = model.merge_and_unload()     # bake the active adapter in for single-use serving

A saved adapter is a folder with adapter_config.json (rank, target modules, base model name) and adapter_model.safetensors (the weights). Version it like code, together with the base model's exact name and revision, the training data version and the eval results.

What it means for operations

  • Storage and versioning. An 84 MB file fits easily in an artefact store; keep every version.
  • Multi-tenant serving. vLLM and similar servers load many LoRA adapters over one base and pick one per request (section 4 covers this).
  • Rollback. Point traffic back to the previous adapter.
  • Base upgrades. A new base model means retraining every adapter. Keep the data and training scripts, not just the adapter files.

A real-life example

A software company offers AI product-description writing to 40 Indian D2C brands, each with its own voice. One adapter per brand at about 84 MB means 3.4 GB of adapters in total, served over one 16 GB base model. Forty fully fine-tuned copies would be 640 GB.

When one brand complains that new descriptions sound "too formal", the team rolls that brand back to last month's adapter in minutes while they investigate. When they later move to a newer base model, they retrain all 40 adapters overnight from stored data — which only works because each adapter's data and config were versioned.

Follow-up questions to expect

  • "Can you use an adapter on a different base model?" — No. The adapter is an update to specific weight matrices; on another model, even a different version of the same family, it is meaningless or harmful.
  • "Merged or unmerged in production?" — Merged for a single task at maximum speed; unmerged when you need many adapters over one base or fast switching.
  • "Can you combine two adapters?" — Yes, with weighted merging, but it needs re-evaluation on both tasks (section 3).