Course Content
Fine-Tuning LLMs
6 sections · 52 lessons
How does a Mixture-of-Experts (MoE) architecture change fine-tuning behavior and cost?
What you need to know
Total versus active parameters
| Model | Total parameters | Active per token |
|---|---|---|
| Mixtral 8x7B | 46.7B | 12.9B |
| Qwen3-30B-A3B (128 experts, 8 used) | 30.5B | 3.3B |
| gpt-oss-120b | 117B | 5.1B |
| DeepSeek-V3 | 671B | 37B |
Compute per token follows the active count, so inference is fast. Memory follows the total count, because every expert must be loaded — any token might need any expert.
What changes when you fine-tune
- Memory is set by total parameters. Qwen3-30B-A3B in bf16 is about 61 GB of weights, like a dense 30B model. In 4-bit it is about 16–17 GB.
- The router can collapse. On a narrow dataset, the router may learn to send most tokens to a few experts. The others stop being used, and you lose capacity. MoE models are trained with an auxiliary load-balancing loss; keep it on during fine-tuning, and freeze the router or give it a very low learning rate.
- Forgetting can be sudden. If some experts drift, the skills they held can drop sharply rather than gradually.
- LoRA placement is a real choice. Attention-only LoRA is cheap and does not touch the router or experts. Adapting the experts means adding LoRA to all of them (many parameters) or some of them (uneven).
- Infrastructure is heavier. Large MoE training needs expert parallelism and all-to-all communication between GPUs, and load imbalance wastes GPU time.
In code (Transformers 5, PEFT 0.2x)
1from transformers import AutoModelForCausalLM2from peft import LoraConfig34model = AutoModelForCausalLM.from_pretrained(5 "Qwen/Qwen3-30B-A3B", dtype="auto",6 output_router_logits=True, # adds the load-balancing loss (coefficient 0.001 by default)7)8lora = LoraConfig(r=16, lora_alpha=32, task_type="CAUSAL_LM",9 target_modules=["q_proj", "k_proj", "v_proj", "o_proj"]) # attention onlyIn Transformers 5, many MoE models store all experts as one fused 3-D weight and the router as a plain parameter, so they are not nn.Linear layers and target_modules does not reach them. To adapt the experts, PEFT has target_parameters. With the attention-only config above, the router and experts stay frozen.
A real-life example
A telecom company chooses Qwen3-30B-A3B for its Hindi customer-support model, because 3B active parameters make serving fast and cheap. The team first estimated fine-tuning memory from the 3B active count and booked a 24 GB GPU. The bf16 model alone needed about 61 GB, so they moved to one 80 GB H100.
The first run trained only on refund conversations, with the load-balancing loss off. Afterwards, routing logs showed most tokens going to a small group of experts, and the model's answers on non-refund topics got noticeably worse. The second run used attention-only LoRA, the router frozen, the load-balancing loss on, and a mix of all ticket types. Refund quality stayed high, and general quality returned to the base model's level. (Made-up details for illustration.)
Follow-up questions to expect
- "Is an MoE cheaper to fine-tune than a dense model?" — Per token it needs less compute than a dense model of the same total size, but the same memory. Compared with a dense model of the same active size, it costs far more memory.
- "Dense 8B or MoE 30B-A3B for a small team?" — The dense 8B is easier to fine-tune and serve on one small GPU. The MoE can give better quality at similar per-token cost if you can afford the memory.
- "What is expert parallelism?" — Placing different experts on different GPUs and sending tokens between GPUs to reach their experts.