Course Content
LLMOps & Deployment
6 sections · 40 lessons
How do distributed computing techniques improve LLM scalability?
What you need to know
The four axes
| Technique | What is split | Communication | Use when |
|---|---|---|---|
| Replication (data parallel) | Nothing — full copies | None between replicas | The model fits on one GPU; you need throughput |
| Tensor parallelism (TP) | Each weight matrix, across GPUs | All-reduce in every layer; needs NVLink | Model too big for one GPU, or lower latency needed |
| Pipeline parallelism (PP) | Groups of layers, across GPUs or nodes | Small, between stages | Model too big for one node |
| Expert parallelism (EP) | MoE experts, across GPUs | All-to-all per MoE layer | Large MoE models |
In vLLM these are flags such as --tensor-parallel-size 2 and --pipeline-parallel-size 2; replicas are separate servers behind a router.
Why "replicate first"
Tensor parallelism adds communication to every layer of every token. Four replicas of a model on one GPU each usually give more total throughput than one copy split across four GPUs, and a failed GPU only takes out a quarter of capacity. TP is the price you pay when the model does not fit.
Disaggregated prefill and decode
Prefill (reading the prompt) is compute-heavy; decode (writing tokens) is memory-bandwidth-heavy. On the same GPU, a long prefill interrupts other users' decoding. Disaggregated serving runs them on separate GPU pools and transfers the KV cache between them, so each pool is tuned for its job. vLLM, SGLang, NVIDIA Dynamo and the llm-d project on Kubernetes support this pattern. It pays off at large scale with long prompts; for small deployments, chunked prefill on one pool is simpler.
Routing across replicas
With many replicas, the router matters: least-outstanding-requests balancing, and prefix- or KV-cache-aware routing so repeated system prompts and conversation turns land where their KV cache already is.
A real-life example
A code assistant for 2,000 engineers moves to a 70B model in FP8 (about 70 GB of weights). One 80 GB H100 would leave only a few GB for KV cache, so each replica uses TP = 2 on an NVLink pair: 35 GB of weights per GPU and about 33 GB of KV cache per GPU — around 216,000 tokens in flight per replica.
A colleague proposes one TP = 8 deployment over all eight GPUs of the node "for maximum speed". Load tests settle it:
- 4 replicas of TP = 2: highest total throughput; losing one GPU removes one replica (25% of capacity) while three keep serving.
- 1 replica of TP = 8: lower time per token for a single request, but lower total throughput because of the extra communication, and one GPU fault takes down everything.
They ship four TP = 2 replicas with KV-cache-aware routing, which raises the prefix-cache hit rate for repeated repository context from 35% to 70%.
Follow-up questions to expect
- "Why keep tensor parallelism inside one node?" — It communicates in every layer; NVLink inside a node is many times faster than the network between nodes.
- "When is pipeline parallelism worth it?" — When a model does not fit on one node's GPUs, such as very large dense or MoE models; it tolerates slower links but adds pipeline bubbles.
- "What does disaggregation cost?" — Moving KV caches between pools needs fast networking and more complex scheduling; it is worth it mainly at large scale.