Course Content
LLMOps & Deployment
6 sections · 40 lessons
What is Mixture of Experts (MoE), and how does it optimize compute usage?
What you need to know
How it works
In each MoE layer, a router scores all experts for the current token and sends it to the top-k (for example 2 of 8, or 8 of 128). Only those experts' feed-forward weights are used. Attention layers are usually shared by all tokens.
Well-known examples
| Model | Total | Active per token |
|---|---|---|
| Mixtral 8x7B | ~47B | ~13B |
| Qwen3-30B-A3B | ~30B | ~3B |
| gpt-oss-120b | ~117B | ~5B |
| DeepSeek-V3 | ~671B | ~37B |
What this means for serving
- Memory follows total. Qwen3-30B-A3B in bf16 needs about 61 GB — like a dense 30B model.
- Speed follows active — at small batch. Each token reads only ~3B parameters of expert and shared weights, so decode can be many times faster than a dense 32B. Rough bandwidth ceilings on an H100: dense 32B (64 GB read per token) about 52 tokens/s; 3B active (about 6.6 GB) about 500 tokens/s.
- At large batch, the gap narrows. Different tokens pick different experts, so more experts are read each step. Compute per token is still much lower.
- Multi-GPU needs expert parallelism — experts spread across GPUs, with all-to-all communication each layer. Uneven routing ("hot" experts) causes load imbalance.
- Quantization helps a lot, because memory is the bottleneck: gpt-oss-120b was released with 4-bit MXFP4 expert weights so it fits on one 80 GB GPU.
A real-life example
A code assistant for 2,000 engineers must pick a self-hosted model for one H100 (80 GB) per replica. The shortlist: a dense 32B coding model and a 30B-total, 3B-active MoE model. Both take about 60–65 GB in bf16, or about half that in FP8.
Load tests (FP8, their real prompts):
- The MoE model streams answers about 3× faster per user at low load and serves more users per GPU at peak.
- The dense 32B model scores 5 points higher on their refactoring eval, but only 1 point higher on explain-and-document tasks.
They deploy both behind the gateway: the MoE model handles explain, document and small edits (70% of traffic) with fast responses, and the dense model handles multi-file refactoring. The router sends requests by task type, and both run in FP8 to leave room for KV cache.
Follow-up questions to expect
- "Why not always use MoE?" — Memory: you pay for all experts. On a small single GPU, or at very low traffic, a dense model of the same memory size is often simpler and good enough.
- "What is load balancing in MoE training?" — An extra training objective (or bias adjustment) that pushes the router to spread tokens across experts, so some are not idle while others are overloaded.
- "How does MoE affect fine-tuning?" — LoRA still works but is usually applied to attention and shared layers or selected experts; check your tools' support.