Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

What is model merging, and how does it pair with fine-tuning strategies?


Scaling the summariser's task vector back toward the base91%64%89%78%83%80%FaithfulnessInstruction followingscale 1.0scale 0.7scale 0.5Illustrative numbers from the lesson's made-up scenario.
Interpolating toward the base trades two points of domain quality for fourteen points of general skill, with no retraining.

What you need to know

The general formula

Text
theta_merged = theta_base + w1 * (theta_1 - theta_base) + w2 * (theta_2 - theta_base) + ...

This works because models fine-tuned from the same base usually stay in the same region of weight space. Averaging them gives a model that still works, which is not true for two models trained from different random starts.

The main methods

MethodWhat it doesWhen to use
Linear / model soupWeighted average of whole modelsCheckpoints or seeds of the same run
SLERPInterpolates along the arc between two models, keeping weight size steadyExactly two models
Task arithmeticAdds scaled task vectors to the baseTwo or more specialists
TIESDrops the smallest changes, picks a sign per weight by vote, averages only the agreeing changesThree or more specialists that interfere
DARERandomly drops a share of each task vector's changes and rescales the rest, then mergesOften combined with TIES (dare_ties)

Interference is the core problem. If one fine-tune pushes a weight up and another pushes it down, a plain average cancels both. TIES and DARE exist to reduce this.

Doing it with mergekit

YAML
# merge.yml — mergekit 0.1.x, run with: mergekit-yaml merge.yml ./mergedmerge_method: tiesbase_model: meta-llama/Llama-3.1-8B-Instructmodels:  - model: ./support-hi-merged       # LoRA already merged into full weights    parameters: {weight: 0.6, density: 0.5}  - model: ./brand-voice-merged    parameters: {weight: 0.4, density: 0.5}parameters:  normalize: truedtype: bfloat16

density: 0.5 keeps the largest half of each task vector's changes. weight sets how much each specialist counts. LoRA fine-tunes need to be merged into full weights first, or merged at the adapter level with PEFT (see the Section 3 lesson on combining LoRA adapters).

How it pairs with fine-tuning

  • Combine specialists — one deployable model instead of a router and several models.
  • Recover from forgetting — interpolate back toward the base: theta = theta_base + 0.7 * (theta_ft - theta_base). You trade a little domain gain for general skills. This idea is known as WiSE-FT.
  • Average checkpoints — the last few checkpoints of one run, or runs with different seeds. Meta's Llama 3 report describes averaging models from different data and hyperparameter runs at the SFT and DPO stages.

A real-life example

A hospital chain fine-tunes an 8B model as a medical-report summariser. Summary quality is good, but the model has started ignoring format instructions from doctors ("in three bullet points"). Their instruction-following eval dropped from 82% (base) to 64% (fine-tune) in their tests.

Instead of retraining, they interpolate between base and fine-tune at three strengths and run both evals:

Task-vector scaleSummary faithfulness (doctor-rated)Instruction following
1.0 (full fine-tune)91%64%
0.789%78%
0.583%80%

They ship 0.7: two points of faithfulness for fourteen points of instruction following, and the whole experiment took an afternoon on one GPU. (Numbers are from this made-up scenario, not a benchmark.)

Follow-up questions to expect

  • "Can you merge a Llama model with a Qwen model?" — No. They have different architectures, tokenizers and no shared base, so their weights do not line up.
  • "Merge or route between models?" — A merge is one artifact with no routing, but tasks can interfere. Routing keeps each model intact but costs serving complexity. Merge when tasks are compatible; route when they conflict.
  • "Does merging replace multi-task training?" — No. Training on a mixed dataset usually beats merging when you can afford it. Merging is cheap and fast, and it is a good first try.