Course Content
Fine-Tuning LLMs
6 sections · 52 lessons
What is model merging, and how does it pair with fine-tuning strategies?
What you need to know
The general formula
theta_merged = theta_base + w1 * (theta_1 - theta_base) + w2 * (theta_2 - theta_base) + ...This works because models fine-tuned from the same base usually stay in the same region of weight space. Averaging them gives a model that still works, which is not true for two models trained from different random starts.
The main methods
| Method | What it does | When to use |
|---|---|---|
| Linear / model soup | Weighted average of whole models | Checkpoints or seeds of the same run |
| SLERP | Interpolates along the arc between two models, keeping weight size steady | Exactly two models |
| Task arithmetic | Adds scaled task vectors to the base | Two or more specialists |
| TIES | Drops the smallest changes, picks a sign per weight by vote, averages only the agreeing changes | Three or more specialists that interfere |
| DARE | Randomly drops a share of each task vector's changes and rescales the rest, then merges | Often combined with TIES (dare_ties) |
Interference is the core problem. If one fine-tune pushes a weight up and another pushes it down, a plain average cancels both. TIES and DARE exist to reduce this.
Doing it with mergekit
1# merge.yml — mergekit 0.1.x, run with: mergekit-yaml merge.yml ./merged2merge_method: ties3base_model: meta-llama/Llama-3.1-8B-Instruct4models:5 - model: ./support-hi-merged # LoRA already merged into full weights6 parameters: {weight: 0.6, density: 0.5}7 - model: ./brand-voice-merged8 parameters: {weight: 0.4, density: 0.5}9parameters:10 normalize: true11dtype: bfloat16density: 0.5 keeps the largest half of each task vector's changes. weight sets how much each specialist counts. LoRA fine-tunes need to be merged into full weights first, or merged at the adapter level with PEFT (see the Section 3 lesson on combining LoRA adapters).
How it pairs with fine-tuning
- Combine specialists — one deployable model instead of a router and several models.
- Recover from forgetting — interpolate back toward the base:
theta = theta_base + 0.7 * (theta_ft - theta_base). You trade a little domain gain for general skills. This idea is known as WiSE-FT. - Average checkpoints — the last few checkpoints of one run, or runs with different seeds. Meta's Llama 3 report describes averaging models from different data and hyperparameter runs at the SFT and DPO stages.
A real-life example
A hospital chain fine-tunes an 8B model as a medical-report summariser. Summary quality is good, but the model has started ignoring format instructions from doctors ("in three bullet points"). Their instruction-following eval dropped from 82% (base) to 64% (fine-tune) in their tests.
Instead of retraining, they interpolate between base and fine-tune at three strengths and run both evals:
| Task-vector scale | Summary faithfulness (doctor-rated) | Instruction following |
|---|---|---|
| 1.0 (full fine-tune) | 91% | 64% |
| 0.7 | 89% | 78% |
| 0.5 | 83% | 80% |
They ship 0.7: two points of faithfulness for fourteen points of instruction following, and the whole experiment took an afternoon on one GPU. (Numbers are from this made-up scenario, not a benchmark.)
Follow-up questions to expect
- "Can you merge a Llama model with a Qwen model?" — No. They have different architectures, tokenizers and no shared base, so their weights do not line up.
- "Merge or route between models?" — A merge is one artifact with no routing, but tasks can interfere. Routing keeps each model intact but costs serving complexity. Merge when tasks are compatible; route when they conflict.
- "Does merging replace multi-task training?" — No. Training on a mixed dataset usually beats merging when you can afford it. Merging is cheap and fast, and it is a good first try.