Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

One enterprise customer sends prompts 50× larger than everyone else, causing noisy-neighbor issues on shared inference servers. How do you isolate workloads and maintain fairness in multi-tenant LLM serving?


What you need to know

Prefill, decode and the KV cache

A 100,000-token prompt has a long prefill that, in a naive server, blocks the batch: every other user's stream pauses for seconds. It also holds a huge KV cache, which can push other requests out of memory. When that happens the server preempts them — pauses them and later recomputes their cache — which wastes more GPU time.

The fairness toolkit

ControlWhat it does
Token buckets per tenantLimits input tokens/sec and output tokens/sec per plan
Per-tenant queues with weighted fair queueingOne tenant's backlog delays only that tenant
Chunked prefillSplits a long prefill into pieces mixed with other requests' decode steps
KV-cache budget per tenantOne tenant cannot evict everyone else
Context caps per tierA clear limit instead of slowing everyone

Chunked prefill is supported by vLLM and other modern engines. On its own it removes most of the visible stutter.

A token bucket per tenant looks like this:

Python
import timeclass TokenBucket:    def __init__(self, rate_per_sec: float, burst: float):        self.rate, self.capacity = rate_per_sec, burst        self.tokens, self.updated = burst, time.monotonic()    def try_take(self, n: int) -> bool:        now = time.monotonic()        self.tokens = min(self.capacity, self.tokens + (now - self.updated) * self.rate)        self.updated = now        if self.tokens >= n:            self.tokens -= n            return True        return Falsebuckets = {"acme": TokenBucket(rate_per_sec=20_000, burst=200_000)}# admit only if buckets[tenant].try_take(prompt_tokens); otherwise queue or return 429

The bucket refills at the tenant's paid rate and allows short bursts. Charging prompt_tokens, not "one request", is what makes a 100K-token prompt count 200 times more than a 500-token one.

When to stop sharing

  1. Meter in tokens — input and output separately.
  2. Queue per tenant — weighted fair scheduling across queues.
  3. Chunk prefill — keep decode flowing for everyone.
  4. Budget KV cache — cap each tenant's share.
  5. Move the outlier — a dedicated replica pool priced for it, or an async batch endpoint where latency does not matter.

The fairness test is simple: the smallest tenant's p95 TTFT should not move when the largest tenant bursts.

A real-life example

Scenario, numbers made up. A legal-tech customer on a shared cluster sends 120K-token contracts, 50 times the average prompt. Each time they run a batch, other tenants' p95 TTFT rises from 0.8 to 6 seconds and streams freeze mid-sentence. The limit was 600 requests per minute per tenant, which the legal-tech customer never reached.

The team turns on chunked prefill, switches to input-token buckets, and caps any tenant at 30% of KV-cache memory. During the next burst, other tenants' p95 TTFT rises only to 1.1 seconds. Then the legal-tech customer's overnight contract review moves to a batch endpoint, and their interactive traffic moves to two dedicated replicas on a higher-priced plan, which they accept because their own latency also improves.

Follow-up questions to expect

  • "What is prefill-decode disaggregation?" — Running prefill and decode on separate GPU pools, so long prefills never interrupt decoding. vLLM, SGLang and NVIDIA Dynamo support it; it helps most when prompts are long and varied.
  • "How do you price the heavy tenant?" — By tokens, with input and output priced separately, and a premium for dedicated capacity.
  • "Isn't a hard context cap unfriendly?" — It is honest. A clear limit per tier is better than every customer getting slower without knowing why.