Transformer Architecture Q&A

Course Content

Transformer Architecture Q&A

6 sections · 60 lessons

What separates dense Transformers from sparse (MoE) models in architecture and compute?


What you need to know

The one architectural difference

Both use the same attention, norms and residual stream. The difference is the FFN: dense models have one per layer; MoE models have a router plus E experts, of which k run per token.

DenseSparse / MoE
Parameters used per tokenalla small fraction (often 5–25%)
FLOPs per token∝ total parameters∝ active parameters
Weight memory∝ total parameters∝ total parameters (all experts resident)
Quality per training FLOPbaselinebetter
Quality per GB of memorybetterworse
Servingsimpleneeds expert parallelism, large batches

Compute: the 6·N·D rule

Training FLOPs for a Transformer are about 6 × N × D, where N is parameters used per token and D is training tokens (2 for the forward pass, 4 for the backward). For MoE, N is the active count.

Text
DeepSeek-V3: 671B total, 37B active, 14.8T tokens (from its technical report)  MoE:    6 × 37e9  × 14.8e12 ≈ 3.3e24 FLOPs  dense:  6 × 671e9 × 14.8e12 ≈ 6.0e25 FLOPs   (if all 671B were active)  → about 18× less compute for the same total size

The report states the full training took about 2.8M H800 GPU hours. A dense model of the same total size would have been far out of reach at that budget.

Memory and serving

All experts must be in memory because any token may pick any of them. 671B parameters at FP8 is about 671 GB — several GPUs before you store any KV cache. Serving spreads experts across GPUs (expert parallelism), and each token's hidden state travels to the GPUs holding its experts and back, every layer.

Efficiency depends on batch size. At batch 1, each expert's weights are read to process one token: poor arithmetic intensity. At large batches, each expert gets many tokens, and the model runs at close to the cost of its active size.

The 2026 picture

Many of the largest open-weight models are MoE: DeepSeek-V3 and R1, Qwen3-235B-A22B (235B total, 22B active), Llama 4, and gpt-oss (about 117B total, 5.1B active for the larger model). Dense models remain common at small sizes and for on-device or single-GPU use, where memory is the limit — for example Llama 3.x 8B, Gemma 3 and Qwen3's dense sizes.

A real-life example

A team building a chatbot for a million users compares a 70B dense model with an MoE of about 400B total and 17B active (Llama 4 Maverick's published size).

Text
compute per token:  dense 70B      vs  MoE ~17B active   → MoE ~4× cheaperweights at FP8:     dense ~70 GB   vs  MoE ~400 GB

At peak, thousands of conversations run at once, so batches are large and the MoE produces far more tokens per GPU-second. They choose MoE for the public chatbot. For a separate on-premises deployment at a bank, where the budget is a single 80 GB GPU, only the dense model fits — so the same team ships a dense model there.

Follow-up questions to expect

  • "Is an MoE faster than a dense model of the same total size?" — Yes in compute, often by the ratio of total to active parameters. Compared with a dense model of the same active size it is slower, because it has more weights to fetch.
  • "Why do MoEs need load balancing?" — Without it, the router favours a few experts, the others get no gradient, and the model behaves like a much smaller one.
  • "Can you distil an MoE into a dense model?" — Yes; train a dense student on the MoE's outputs. You keep some of the quality at a fraction of the memory.