Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

How does a Mixture-of-Experts (MoE) architecture change fine-tuning behavior and cost?


Dense 8B against Qwen3-30B-A3BDense 8B• All 8B parameters work on every token• About 16 GB of bf16 weights to train• LoRA fits one 24 to 40 GB GPU• No router that can collapseMoE 30B-A3B• 3.3B of 30.5B parameters per token• About 61 GB of bf16 weights to train• Needs an 80 GB GPU, or QLoRA• Router can collapse on narrow data
MoE makes serving cheap per token, but fine-tuning pays for every expert in memory.

What you need to know

Total versus active parameters

ModelTotal parametersActive per token
Mixtral 8x7B46.7B12.9B
Qwen3-30B-A3B (128 experts, 8 used)30.5B3.3B
gpt-oss-120b117B5.1B
DeepSeek-V3671B37B

Compute per token follows the active count, so inference is fast. Memory follows the total count, because every expert must be loaded — any token might need any expert.

What changes when you fine-tune

  • Memory is set by total parameters. Qwen3-30B-A3B in bf16 is about 61 GB of weights, like a dense 30B model. In 4-bit it is about 16–17 GB.
  • The router can collapse. On a narrow dataset, the router may learn to send most tokens to a few experts. The others stop being used, and you lose capacity. MoE models are trained with an auxiliary load-balancing loss; keep it on during fine-tuning, and freeze the router or give it a very low learning rate.
  • Forgetting can be sudden. If some experts drift, the skills they held can drop sharply rather than gradually.
  • LoRA placement is a real choice. Attention-only LoRA is cheap and does not touch the router or experts. Adapting the experts means adding LoRA to all of them (many parameters) or some of them (uneven).
  • Infrastructure is heavier. Large MoE training needs expert parallelism and all-to-all communication between GPUs, and load imbalance wastes GPU time.

In code (Transformers 5, PEFT 0.2x)

Python
from transformers import AutoModelForCausalLMfrom peft import LoraConfigmodel = AutoModelForCausalLM.from_pretrained(    "Qwen/Qwen3-30B-A3B", dtype="auto",    output_router_logits=True,     # adds the load-balancing loss (coefficient 0.001 by default))lora = LoraConfig(r=16, lora_alpha=32, task_type="CAUSAL_LM",                  target_modules=["q_proj", "k_proj", "v_proj", "o_proj"])  # attention only

In Transformers 5, many MoE models store all experts as one fused 3-D weight and the router as a plain parameter, so they are not nn.Linear layers and target_modules does not reach them. To adapt the experts, PEFT has target_parameters. With the attention-only config above, the router and experts stay frozen.

A real-life example

A telecom company chooses Qwen3-30B-A3B for its Hindi customer-support model, because 3B active parameters make serving fast and cheap. The team first estimated fine-tuning memory from the 3B active count and booked a 24 GB GPU. The bf16 model alone needed about 61 GB, so they moved to one 80 GB H100.

The first run trained only on refund conversations, with the load-balancing loss off. Afterwards, routing logs showed most tokens going to a small group of experts, and the model's answers on non-refund topics got noticeably worse. The second run used attention-only LoRA, the router frozen, the load-balancing loss on, and a mix of all ticket types. Refund quality stayed high, and general quality returned to the base model's level. (Made-up details for illustration.)

Follow-up questions to expect

  • "Is an MoE cheaper to fine-tune than a dense model?" — Per token it needs less compute than a dense model of the same total size, but the same memory. Compared with a dense model of the same active size, it costs far more memory.
  • "Dense 8B or MoE 30B-A3B for a small team?" — The dense 8B is easier to fine-tune and serve on one small GPU. The MoE can give better quality at similar per-token cost if you can afford the memory.
  • "What is expert parallelism?" — Placing different experts on different GPUs and sending tokens between GPUs to reach their experts.