Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

How can you combine or merge multiple LoRA adapters?


Clause adapter plus Hindi adapter, three ways-3 pointsholds-1 pointholdsno dropholdsClause accuracyHindi summary qualitycat mergedare_ties, density 0.5One multi-task adapter
Hindi quality held in every variant — only the clause eval exposed the interference, which is why every merged task needs its own eval.

What you need to know

Weight-space merging in PEFT 0.21

Python
model.add_weighted_adapter(    adapters=["clauses", "hindi"], weights=[1.0, 1.0],    adapter_name="clauses_hindi", combination_type="cat")model.set_adapter("clauses_hindi")
combination_typeWhat it doesWatch out for
catStacks the adapters' ranks side by side: the exact sum of the updatesRank grows (16 + 16 = 32), so the adapter grows
svd (default)Adds the full updates, then compresses back to a chosen rank with SVDSlower; the compression loses a little
linearWeighted sum of the A matrices and of the B matrices separatelyNeeds equal ranks; only approximates adding the updates
ties, dare_ties, dare_linear (and _svd versions)Drop small or random entries, resolve sign conflicts, then combineNeeds a density, such as 0.5

TIES (2023) keeps only the largest changes in each adapter, picks a majority sign for each weight, and averages only the values that agree. DARE (2023) randomly drops a share of each update and rescales the rest. Both reduce interference when three or more adapters are merged.

Other options

  • Stacking. Merge adapter A into the base, then train adapter B on top. Predictable, but order matters and you cannot unmix them later.
  • Routing. Keep adapters separate and choose per request — a simple classifier or a rule picks the adapter. No interference at all, at the cost of a routing step.
  • Multi-task training. Train one adapter on both datasets mixed. Usually the best quality when you have the data.

Rules

  • Merge only adapters trained on the identical base checkpoint.
  • Evaluate the merged model on every original task, not just the average.

A real-life example

The Mumbai law firm has two adapters on the same base: one classifies clauses, the other writes plain-Hindi summaries for clients. It wants one model that classifies a clause and explains it in Hindi.

  • cat merge: Hindi quality holds, but clause accuracy drops 3 points — the Hindi adapter's changes interfere.
  • dare_ties with density 0.5: clause accuracy drops only 1 point.
  • Multi-task adapter trained on both datasets: no measurable drop on either task.

The team ships the multi-task adapter, and keeps dare_ties merging as the quick option for experiments.

Follow-up questions to expect

  • "Why does merging hurt a task?" — Two updates can push the same weights in opposite directions, so the sum cancels part of each skill.
  • "Can you subtract an adapter?" — Negative weights are allowed, and "task negation" can reduce a behaviour, but it is unpredictable; test carefully.
  • "How is this different from full-model merging?" — Same ideas (section 4 covers model merging of full checkpoints); adapter merging is cheaper because the updates are small.