LLMs Deep Dive

Course Content

LLMs Deep Dive

10 sections · 40 lessons

How does Mixture of Experts (MoE) improve scalability?


Top-2 routing, four experts shown·0.62·0.380.55·0.45··0.700.30·0.51··0.49E1E2E3E4refundkabmilegaRsIllustrative router weights after renormalising the top two.
Every expert must sit in memory, but each token pays compute for only two — that gap is the whole MoE trade.

What you need to know

Routing one token, with numbers

A layer has 8 experts and uses top-2 routing. For the token " refund" the router's softmax gives:

Text
E1 0.05 | E2 0.40 | E3 0.08 | E4 0.25 | E5 0.07 | E6 0.05 | E7 0.06 | E8 0.04

Top-2 are E2 (0.40) and E4 (0.25). Renormalised: 0.40 / 0.65 ≈ 0.62 and 0.25 / 0.65 ≈ 0.38.

Text
output = 0.62 × E2(x) + 0.38 × E4(x)

The other six experts do no work for this token. Attention layers are usually still dense and shared.

Parameters versus compute — real published sizes

ModelTotal parametersActive per token
Mixtral 8x7Babout 47Babout 13B (2 of 8 experts)
Qwen3-235B-A22B235B22B
DeepSeek-V3671B37B

"Total" decides memory; "active" decides compute per token. Mixtral needs memory for 47B parameters (about 94 GB in 16-bit) but does roughly the arithmetic of a 13B dense model per token.

The costs

  • Memory — all experts must be loaded even though most are idle for any one token.
  • Load balancing — without it, the router sends most tokens to a few favourite experts and the rest never learn. Fixes include an auxiliary balancing loss, a per-expert capacity limit, and DeepSeek-V3's bias-adjustment approach that balances without an auxiliary loss.
  • Communication — experts are spread across GPUs (expert parallelism), so tokens are shuffled between GPUs twice per MoE layer.
  • Batching effects — at batch size 1, only active experts' weights are read, so decoding is fast; at large batch sizes, different tokens hit different experts and most weights are read anyway.
  • Fine-tuning — routers can become unstable and experts overfit; many teams fine-tune only attention and shared parts, or use LoRA.

Where MoE fits in scaling

Scaling laws show loss falls predictably as parameters, data and compute grow. The 2022 Chinchilla study found that, for a fixed training budget, parameters and training tokens should grow together (about 20 tokens per parameter). Since then, models have been trained on far more tokens than that because inference cost matters more than training cost — Meta trained Llama 3 8B on more than 15 trillion tokens. MoE is another lever: more parameters without more compute per token. Reasoning models add a fourth lever, spending more compute at inference time by thinking longer.

A real-life example

An e-commerce company self-hosts a model for its search assistant and compares a dense 14B model with an MoE model of about 30B total and 3B active parameters (the Qwen3-30B-A3B design is one such shape). On their 2,000-query evaluation set, the MoE model is about as accurate as the dense one.

Serving tells the real story. The MoE model generates tokens noticeably faster at low load, because each token reads only about 3B parameters' worth of weights. But it needs memory for 30B parameters (about 60 GB in 16-bit) versus 28 GB for the dense model, so it needs a larger GPU. At peak sale traffic with large batches, the speed advantage shrinks, as most experts are active across the batch. They choose MoE for the latency-sensitive chat and keep the dense model for a batch job that tags products overnight.

Follow-up questions to expect

  • "Do experts specialise in topics like maths or Hindi?" — Studies of trained routers find specialisation is mostly on token types and syntax, not clean human topics.
  • "What is a shared expert?" — An expert every token always uses, alongside routed ones, to hold common knowledge; DeepSeek's MoE design uses them.
  • "Is MoE always better than dense?" — Per unit of training compute it usually is; but when memory is the bottleneck, such as on one small GPU or a phone, a dense model of the same memory size is often the better choice.