Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your LLM API has p50 of 800ms but p99 of 14 seconds. Most users are happy; 1% rage-quit. How do you cut tail latency without changing models?


What you need to know

Total latency for an LLM call is roughly: queue time + prefill (reading the prompt) + output tokens x time per token. That formula tells you where to look.

Diagnose before fixing

Log per request: input tokens, output tokens, time-to-first-token (TTFT) and total time. Then look at the slowest 1%. Usually the causes are:

CauseSignature in the logs
Very long outputsTotal time scales with output tokens; TTFT normal
Retries stackingTwo or three attempts in the trace, each near the timeout
Provider queueing at peakTTFT high, output length normal, clustered at certain hours
Very long promptsTTFT high on requests with huge context

Fixes, in the order I'd apply them

  1. Cap max_tokens — to what the feature needs. Generation time is roughly linear in output length; an uncapped request can run 30 times longer than a typical one.
  2. Hedge — if no first token by about the p95 TTFT, send a second identical request and use whichever answers first; cancel the other.
  3. Fix timeouts and retries — a 30-second timeout plus one retry means a 60-second worst case. Use a short first-token timeout and a total budget across all attempts.
  4. Stream — users perceive TTFT, not total time.
  5. Shrink the prompt — prefill grows with input; trim history and retrieved context.
  6. Fall back — past the budget, answer from a faster tier or a cached response.
Python
async def hedged(call, hedge_after=2.0):    first = asyncio.create_task(call())    done, _ = await asyncio.wait({first}, timeout=hedge_after)    if done:        return first.result()    second = asyncio.create_task(call())    done, pending = await asyncio.wait({first, second}, return_when=asyncio.FIRST_COMPLETED)    for p in pending:        p.cancel()    return done.pop().result()

In production you hedge on time-to-first-token of a streaming call, so the duplicate fires only when the first request has not started answering.

Manage by the right metric

Track p99 TTFT and p99 total latency per route. An average hides the tail completely.

A real-life example

Scenario (illustrative numbers). A B2B analytics product's "explain this chart" feature has p50 800 ms and p99 14 s. Logs show two groups in the slowest 1%: 60% are answers of 1,500 to 3,000 tokens where 300 would do, because max_tokens was left at 4,096; 30% are requests with a retry after a 6-second timeout at peak hours.

The team caps output at 500 tokens and tells the model to be brief, adds hedging at 2 seconds of no first token, and replaces timeout-plus-retry with a single 8-second budget. p99 falls to 3.1 s, and hedging fires on 3% of requests, adding about 3% to spend. Streaming makes the perceived wait under a second for most users.

Follow-up questions to expect

  • "Doesn't hedging double the cost?" — Only for hedged requests. If it fires on 3% of traffic, cost rises about 3%; cancel the loser as soon as one wins.
  • "What if the provider itself is slow for everyone?" — Then it is an incident, not a tail; the circuit breaker should move traffic to a fallback model or region.
  • "Why not just use a smaller, faster model?" — The question rules it out, and usually the tail comes from request shape, not model speed.