Course Content
Transformer Architecture Q&A
6 sections · 60 lessons
What separates dense Transformers from sparse (MoE) models in architecture and compute?
What you need to know
The one architectural difference
Both use the same attention, norms and residual stream. The difference is the FFN: dense models have one per layer; MoE models have a router plus E experts, of which k run per token.
| Dense | Sparse / MoE | |
|---|---|---|
| Parameters used per token | all | a small fraction (often 5–25%) |
| FLOPs per token | ∝ total parameters | ∝ active parameters |
| Weight memory | ∝ total parameters | ∝ total parameters (all experts resident) |
| Quality per training FLOP | baseline | better |
| Quality per GB of memory | better | worse |
| Serving | simple | needs expert parallelism, large batches |
Compute: the 6·N·D rule
Training FLOPs for a Transformer are about 6 × N × D, where N is parameters used per token and D is training tokens (2 for the forward pass, 4 for the backward). For MoE, N is the active count.
DeepSeek-V3: 671B total, 37B active, 14.8T tokens (from its technical report) MoE: 6 × 37e9 × 14.8e12 ≈ 3.3e24 FLOPs dense: 6 × 671e9 × 14.8e12 ≈ 6.0e25 FLOPs (if all 671B were active) → about 18× less compute for the same total sizeThe report states the full training took about 2.8M H800 GPU hours. A dense model of the same total size would have been far out of reach at that budget.
Memory and serving
All experts must be in memory because any token may pick any of them. 671B parameters at FP8 is about 671 GB — several GPUs before you store any KV cache. Serving spreads experts across GPUs (expert parallelism), and each token's hidden state travels to the GPUs holding its experts and back, every layer.
Efficiency depends on batch size. At batch 1, each expert's weights are read to process one token: poor arithmetic intensity. At large batches, each expert gets many tokens, and the model runs at close to the cost of its active size.
The 2026 picture
Many of the largest open-weight models are MoE: DeepSeek-V3 and R1, Qwen3-235B-A22B (235B total, 22B active), Llama 4, and gpt-oss (about 117B total, 5.1B active for the larger model). Dense models remain common at small sizes and for on-device or single-GPU use, where memory is the limit — for example Llama 3.x 8B, Gemma 3 and Qwen3's dense sizes.
A real-life example
A team building a chatbot for a million users compares a 70B dense model with an MoE of about 400B total and 17B active (Llama 4 Maverick's published size).
compute per token: dense 70B vs MoE ~17B active → MoE ~4× cheaperweights at FP8: dense ~70 GB vs MoE ~400 GBAt peak, thousands of conversations run at once, so batches are large and the MoE produces far more tokens per GPU-second. They choose MoE for the public chatbot. For a separate on-premises deployment at a bank, where the budget is a single 80 GB GPU, only the dense model fits — so the same team ships a dense model there.
Follow-up questions to expect
- "Is an MoE faster than a dense model of the same total size?" — Yes in compute, often by the ratio of total to active parameters. Compared with a dense model of the same active size it is slower, because it has more weights to fetch.
- "Why do MoEs need load balancing?" — Without it, the router favours a few experts, the others get no gradient, and the model behaves like a much smaller one.
- "Can you distil an MoE into a dense model?" — Yes; train a dense student on the MoE's outputs. You keep some of the quality at a fraction of the memory.