Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
One enterprise customer sends prompts 50× larger than everyone else, causing noisy-neighbor issues on shared inference servers. How do you isolate workloads and maintain fairness in multi-tenant LLM serving?
What you need to know
Prefill, decode and the KV cache
A 100,000-token prompt has a long prefill that, in a naive server, blocks the batch: every other user's stream pauses for seconds. It also holds a huge KV cache, which can push other requests out of memory. When that happens the server preempts them — pauses them and later recomputes their cache — which wastes more GPU time.
The fairness toolkit
| Control | What it does |
|---|---|
| Token buckets per tenant | Limits input tokens/sec and output tokens/sec per plan |
| Per-tenant queues with weighted fair queueing | One tenant's backlog delays only that tenant |
| Chunked prefill | Splits a long prefill into pieces mixed with other requests' decode steps |
| KV-cache budget per tenant | One tenant cannot evict everyone else |
| Context caps per tier | A clear limit instead of slowing everyone |
Chunked prefill is supported by vLLM and other modern engines. On its own it removes most of the visible stutter.
A token bucket per tenant looks like this:
1import time23class TokenBucket:4 def __init__(self, rate_per_sec: float, burst: float):5 self.rate, self.capacity = rate_per_sec, burst6 self.tokens, self.updated = burst, time.monotonic()78 def try_take(self, n: int) -> bool:9 now = time.monotonic()10 self.tokens = min(self.capacity, self.tokens + (now - self.updated) * self.rate)11 self.updated = now12 if self.tokens >= n:13 self.tokens -= n14 return True15 return False1617buckets = {"acme": TokenBucket(rate_per_sec=20_000, burst=200_000)}18# admit only if buckets[tenant].try_take(prompt_tokens); otherwise queue or return 429The bucket refills at the tenant's paid rate and allows short bursts. Charging prompt_tokens, not "one request", is what makes a 100K-token prompt count 200 times more than a 500-token one.
When to stop sharing
- Meter in tokens — input and output separately.
- Queue per tenant — weighted fair scheduling across queues.
- Chunk prefill — keep decode flowing for everyone.
- Budget KV cache — cap each tenant's share.
- Move the outlier — a dedicated replica pool priced for it, or an async batch endpoint where latency does not matter.
The fairness test is simple: the smallest tenant's p95 TTFT should not move when the largest tenant bursts.
A real-life example
Scenario, numbers made up. A legal-tech customer on a shared cluster sends 120K-token contracts, 50 times the average prompt. Each time they run a batch, other tenants' p95 TTFT rises from 0.8 to 6 seconds and streams freeze mid-sentence. The limit was 600 requests per minute per tenant, which the legal-tech customer never reached.
The team turns on chunked prefill, switches to input-token buckets, and caps any tenant at 30% of KV-cache memory. During the next burst, other tenants' p95 TTFT rises only to 1.1 seconds. Then the legal-tech customer's overnight contract review moves to a batch endpoint, and their interactive traffic moves to two dedicated replicas on a higher-priced plan, which they accept because their own latency also improves.
Follow-up questions to expect
- "What is prefill-decode disaggregation?" — Running prefill and decode on separate GPU pools, so long prefills never interrupt decoding. vLLM, SGLang and NVIDIA Dynamo support it; it helps most when prompts are long and varied.
- "How do you price the heavy tenant?" — By tokens, with input and output priced separately, and a premium for dedicated capacity.
- "Isn't a hard context cap unfriendly?" — It is honest. A clear limit per tier is better than every customer getting slower without knowing why.