Course Content
LLMOps & Deployment
6 sections · 40 lessons
What are the key performance metrics for LLM systems (latency, throughput, availability)?
What you need to know
Two phases of generation
- Prefill — the model reads the whole prompt in one parallel pass and builds the KV cache. Compute-heavy. Its time grows with input length. It sets TTFT.
- Decode — the model makes one token per step. Each step reads all the model weights from GPU memory, so it is limited by memory bandwidth. It sets TPOT (also called inter-token latency, ITL).
The metrics
| Family | Metric | Why |
|---|---|---|
| Latency | TTFT | What users feel first; grows with queue and prompt length |
| Latency | TPOT / ITL | 20–50 ms per token is faster than reading speed |
| Latency | End-to-end ≈ TTFT + output_tokens × TPOT | What a non-streaming caller waits |
| Throughput | Output tokens/s per GPU | Sets cost per million tokens |
| Throughput | Goodput | Requests per second that also met the latency target |
| Saturation | Queue depth, KV-cache usage | Rise tens of seconds before latency spikes |
| Availability | Success rate by 5xx / 429 / timeout; fallback rate | A degraded answer is not a full success |
On vLLM these appear as Prometheus metrics such as vllm:time_to_first_token_seconds, vllm:inter_token_latency_seconds, vllm:num_requests_waiting and vllm:kv_cache_usage_perc.
The batch-size trade-off
A decode step reads all weights once, whether it serves 1 sequence or 64. So batching raises total throughput a lot — but each step gets a bit slower, so every user sees slower tokens and new requests may wait longer. Illustrative numbers for an 8B model on one H100:
| Batch size | TPOT | Tokens/s per user | Tokens/s per GPU |
|---|---|---|---|
| 1 | 7 ms | 143 | 143 |
| 8 | 8 ms | 125 | 1,000 |
| 32 | 12 ms | 83 | 2,670 |
| 64 | 18 ms | 56 | 3,560 |
Going from batch 1 to 64 makes the GPU 25× more productive, while each user's speed falls from 143 to 56 tokens a second — still faster than reading. That is why continuous batching is the default. The limit is your TPOT target and KV-cache memory.
Reporting rules
- Percentiles, not averages. A 2-second mean can hide a 20-second p99.
- Always next to token counts. "p95 = 6 s" means nothing without "at 3,000 input and 400 output tokens".
A real-life example
A fintech support bot sets these SLOs: TTFT p95 under 1.0 s, TPOT p95 under 50 ms, success rate 99.5%, fallback rate under 2%.
During the Diwali sale, TTFT p95 climbs from 0.6 s to 2.8 s while TPOT stays at 30 ms. The split tells the story: decode is fine, so the problem is before the first token. The dashboard shows num_requests_waiting rising from 0 to 180 fifteen minutes earlier — a queueing problem, not a model problem. The fix is scaling on queue depth, not changing the model. The next sale, an alert on queue depth fires before users notice.
Follow-up questions to expect
- "What drives TTFT up?" — Queueing, long prompts (prefill), cold starts, and prefill of other requests competing in the same batch.
- "How do you raise throughput without hurting latency?" — Continuous batching, prefix caching, quantization, and a TPOT target that caps batch size.
- "What is goodput?" — Throughput counting only requests that met the SLO; high raw throughput with blown latency is not useful work.