Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your LLM API has p50 of 800ms but p99 of 14 seconds. Most users are happy; 1% rage-quit. How do you cut tail latency without changing models?
What you need to know
Total latency for an LLM call is roughly: queue time + prefill (reading the prompt) + output tokens x time per token. That formula tells you where to look.
Diagnose before fixing
Log per request: input tokens, output tokens, time-to-first-token (TTFT) and total time. Then look at the slowest 1%. Usually the causes are:
| Cause | Signature in the logs |
|---|---|
| Very long outputs | Total time scales with output tokens; TTFT normal |
| Retries stacking | Two or three attempts in the trace, each near the timeout |
| Provider queueing at peak | TTFT high, output length normal, clustered at certain hours |
| Very long prompts | TTFT high on requests with huge context |
Fixes, in the order I'd apply them
- Cap
max_tokens— to what the feature needs. Generation time is roughly linear in output length; an uncapped request can run 30 times longer than a typical one. - Hedge — if no first token by about the p95 TTFT, send a second identical request and use whichever answers first; cancel the other.
- Fix timeouts and retries — a 30-second timeout plus one retry means a 60-second worst case. Use a short first-token timeout and a total budget across all attempts.
- Stream — users perceive TTFT, not total time.
- Shrink the prompt — prefill grows with input; trim history and retrieved context.
- Fall back — past the budget, answer from a faster tier or a cached response.
1async def hedged(call, hedge_after=2.0):2 first = asyncio.create_task(call())3 done, _ = await asyncio.wait({first}, timeout=hedge_after)4 if done:5 return first.result()6 second = asyncio.create_task(call())7 done, pending = await asyncio.wait({first, second}, return_when=asyncio.FIRST_COMPLETED)8 for p in pending:9 p.cancel()10 return done.pop().result()In production you hedge on time-to-first-token of a streaming call, so the duplicate fires only when the first request has not started answering.
Manage by the right metric
Track p99 TTFT and p99 total latency per route. An average hides the tail completely.
A real-life example
Scenario (illustrative numbers). A B2B analytics product's "explain this chart" feature has p50 800 ms and p99 14 s. Logs show two groups in the slowest 1%: 60% are answers of 1,500 to 3,000 tokens where 300 would do, because max_tokens was left at 4,096; 30% are requests with a retry after a 6-second timeout at peak hours.
The team caps output at 500 tokens and tells the model to be brief, adds hedging at 2 seconds of no first token, and replaces timeout-plus-retry with a single 8-second budget. p99 falls to 3.1 s, and hedging fires on 3% of requests, adding about 3% to spend. Streaming makes the perceived wait under a second for most users.
Follow-up questions to expect
- "Doesn't hedging double the cost?" — Only for hedged requests. If it fires on 3% of traffic, cost rises about 3%; cancel the loser as soon as one wins.
- "What if the provider itself is slow for everyone?" — Then it is an incident, not a tail; the circuit breaker should move traffic to a fallback model or region.
- "Why not just use a smaller, faster model?" — The question rules it out, and usually the tail comes from request shape, not model speed.