LLMOps & Deployment

Course Content

LLMOps & Deployment

6 sections · 40 lessons

What is Mixture of Experts (MoE), and how does it optimize compute usage?


Dense 32B against a 30B MoE with 3B activeDense 32B• About 64 GB in bf16• Reads 64 GB of weights per token• Bandwidth ceiling near 52 tokens/s• Won by 5 points on refactoringMoE 30B, 3B active• About 61 GB in bf16 — all experts• Reads about 6.6 GB per token• Bandwidth ceiling near 500 tokens/s• Tied on explain and document tasks
Memory follows total parameters and speed follows active ones, so the team kept both and routed by task type.

What you need to know

How it works

In each MoE layer, a router scores all experts for the current token and sends it to the top-k (for example 2 of 8, or 8 of 128). Only those experts' feed-forward weights are used. Attention layers are usually shared by all tokens.

Well-known examples

ModelTotalActive per token
Mixtral 8x7B~47B~13B
Qwen3-30B-A3B~30B~3B
gpt-oss-120b~117B~5B
DeepSeek-V3~671B~37B

What this means for serving

  • Memory follows total. Qwen3-30B-A3B in bf16 needs about 61 GB — like a dense 30B model.
  • Speed follows active — at small batch. Each token reads only ~3B parameters of expert and shared weights, so decode can be many times faster than a dense 32B. Rough bandwidth ceilings on an H100: dense 32B (64 GB read per token) about 52 tokens/s; 3B active (about 6.6 GB) about 500 tokens/s.
  • At large batch, the gap narrows. Different tokens pick different experts, so more experts are read each step. Compute per token is still much lower.
  • Multi-GPU needs expert parallelism — experts spread across GPUs, with all-to-all communication each layer. Uneven routing ("hot" experts) causes load imbalance.
  • Quantization helps a lot, because memory is the bottleneck: gpt-oss-120b was released with 4-bit MXFP4 expert weights so it fits on one 80 GB GPU.

A real-life example

A code assistant for 2,000 engineers must pick a self-hosted model for one H100 (80 GB) per replica. The shortlist: a dense 32B coding model and a 30B-total, 3B-active MoE model. Both take about 60–65 GB in bf16, or about half that in FP8.

Load tests (FP8, their real prompts):

  • The MoE model streams answers about 3× faster per user at low load and serves more users per GPU at peak.
  • The dense 32B model scores 5 points higher on their refactoring eval, but only 1 point higher on explain-and-document tasks.

They deploy both behind the gateway: the MoE model handles explain, document and small edits (70% of traffic) with fast responses, and the dense model handles multi-file refactoring. The router sends requests by task type, and both run in FP8 to leave room for KV cache.

Follow-up questions to expect

  • "Why not always use MoE?" — Memory: you pay for all experts. On a small single GPU, or at very low traffic, a dense model of the same memory size is often simpler and good enough.
  • "What is load balancing in MoE training?" — An extra training objective (or bias adjustment) that pushes the router to spread tokens across experts, so some are not idle while others are overloaded.
  • "How does MoE affect fine-tuning?" — LoRA still works but is usually applied to attention and shared layers or selected experts; check your tools' support.