Course Content
LLMs Deep Dive
10 sections · 40 lessons
How does Mixture of Experts (MoE) improve scalability?
What you need to know
Routing one token, with numbers
A layer has 8 experts and uses top-2 routing. For the token " refund" the router's softmax gives:
E1 0.05 | E2 0.40 | E3 0.08 | E4 0.25 | E5 0.07 | E6 0.05 | E7 0.06 | E8 0.04Top-2 are E2 (0.40) and E4 (0.25). Renormalised: 0.40 / 0.65 ≈ 0.62 and 0.25 / 0.65 ≈ 0.38.
output = 0.62 × E2(x) + 0.38 × E4(x)The other six experts do no work for this token. Attention layers are usually still dense and shared.
Parameters versus compute — real published sizes
| Model | Total parameters | Active per token |
|---|---|---|
| Mixtral 8x7B | about 47B | about 13B (2 of 8 experts) |
| Qwen3-235B-A22B | 235B | 22B |
| DeepSeek-V3 | 671B | 37B |
"Total" decides memory; "active" decides compute per token. Mixtral needs memory for 47B parameters (about 94 GB in 16-bit) but does roughly the arithmetic of a 13B dense model per token.
The costs
- Memory — all experts must be loaded even though most are idle for any one token.
- Load balancing — without it, the router sends most tokens to a few favourite experts and the rest never learn. Fixes include an auxiliary balancing loss, a per-expert capacity limit, and DeepSeek-V3's bias-adjustment approach that balances without an auxiliary loss.
- Communication — experts are spread across GPUs (expert parallelism), so tokens are shuffled between GPUs twice per MoE layer.
- Batching effects — at batch size 1, only active experts' weights are read, so decoding is fast; at large batch sizes, different tokens hit different experts and most weights are read anyway.
- Fine-tuning — routers can become unstable and experts overfit; many teams fine-tune only attention and shared parts, or use LoRA.
Where MoE fits in scaling
Scaling laws show loss falls predictably as parameters, data and compute grow. The 2022 Chinchilla study found that, for a fixed training budget, parameters and training tokens should grow together (about 20 tokens per parameter). Since then, models have been trained on far more tokens than that because inference cost matters more than training cost — Meta trained Llama 3 8B on more than 15 trillion tokens. MoE is another lever: more parameters without more compute per token. Reasoning models add a fourth lever, spending more compute at inference time by thinking longer.
A real-life example
An e-commerce company self-hosts a model for its search assistant and compares a dense 14B model with an MoE model of about 30B total and 3B active parameters (the Qwen3-30B-A3B design is one such shape). On their 2,000-query evaluation set, the MoE model is about as accurate as the dense one.
Serving tells the real story. The MoE model generates tokens noticeably faster at low load, because each token reads only about 3B parameters' worth of weights. But it needs memory for 30B parameters (about 60 GB in 16-bit) versus 28 GB for the dense model, so it needs a larger GPU. At peak sale traffic with large batches, the speed advantage shrinks, as most experts are active across the batch. They choose MoE for the latency-sensitive chat and keep the dense model for a batch job that tags products overnight.
Follow-up questions to expect
- "Do experts specialise in topics like maths or Hindi?" — Studies of trained routers find specialisation is mostly on token types and syntax, not clean human topics.
- "What is a shared expert?" — An expert every token always uses, alongside routed ones, to hold common knowledge; DeepSeek's MoE design uses them.
- "Is MoE always better than dense?" — Per unit of training compute it usually is; but when memory is the bottleneck, such as on one small GPU or a phone, a dense model of the same memory size is often the better choice.