LLMOps & Deployment

Course Content

LLMOps & Deployment

6 sections · 40 lessons

What are the key performance metrics for LLM systems (latency, throughput, availability)?


Batch size against speed: 8B model on one H100 (illustrative)714314381251,00012832,67018563,560TPOT mstok/s per usertok/s per GPUbatch 1batch 8batch 32batch 64Each decode step reads the weights once, however many sequences share it.
Batching makes the GPU 25 times more productive while each user still reads faster than they can — the latency target decides where to stop.

What you need to know

Two phases of generation

  • Prefill — the model reads the whole prompt in one parallel pass and builds the KV cache. Compute-heavy. Its time grows with input length. It sets TTFT.
  • Decode — the model makes one token per step. Each step reads all the model weights from GPU memory, so it is limited by memory bandwidth. It sets TPOT (also called inter-token latency, ITL).

The metrics

FamilyMetricWhy
LatencyTTFTWhat users feel first; grows with queue and prompt length
LatencyTPOT / ITL20–50 ms per token is faster than reading speed
LatencyEnd-to-end ≈ TTFT + output_tokens × TPOTWhat a non-streaming caller waits
ThroughputOutput tokens/s per GPUSets cost per million tokens
ThroughputGoodputRequests per second that also met the latency target
SaturationQueue depth, KV-cache usageRise tens of seconds before latency spikes
AvailabilitySuccess rate by 5xx / 429 / timeout; fallback rateA degraded answer is not a full success

On vLLM these appear as Prometheus metrics such as vllm:time_to_first_token_seconds, vllm:inter_token_latency_seconds, vllm:num_requests_waiting and vllm:kv_cache_usage_perc.

The batch-size trade-off

A decode step reads all weights once, whether it serves 1 sequence or 64. So batching raises total throughput a lot — but each step gets a bit slower, so every user sees slower tokens and new requests may wait longer. Illustrative numbers for an 8B model on one H100:

Batch sizeTPOTTokens/s per userTokens/s per GPU
17 ms143143
88 ms1251,000
3212 ms832,670
6418 ms563,560

Going from batch 1 to 64 makes the GPU 25× more productive, while each user's speed falls from 143 to 56 tokens a second — still faster than reading. That is why continuous batching is the default. The limit is your TPOT target and KV-cache memory.

Reporting rules

  • Percentiles, not averages. A 2-second mean can hide a 20-second p99.
  • Always next to token counts. "p95 = 6 s" means nothing without "at 3,000 input and 400 output tokens".

A real-life example

A fintech support bot sets these SLOs: TTFT p95 under 1.0 s, TPOT p95 under 50 ms, success rate 99.5%, fallback rate under 2%.

During the Diwali sale, TTFT p95 climbs from 0.6 s to 2.8 s while TPOT stays at 30 ms. The split tells the story: decode is fine, so the problem is before the first token. The dashboard shows num_requests_waiting rising from 0 to 180 fifteen minutes earlier — a queueing problem, not a model problem. The fix is scaling on queue depth, not changing the model. The next sale, an alert on queue depth fires before users notice.

Follow-up questions to expect

  • "What drives TTFT up?" — Queueing, long prompts (prefill), cold starts, and prefill of other requests competing in the same batch.
  • "How do you raise throughput without hurting latency?" — Continuous batching, prefix caching, quantization, and a TPOT target that caps batch size.
  • "What is goodput?" — Throughput counting only requests that met the SLO; high raw throughput with blown latency is not useful work.