Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your AI product goes viral overnight. GPU utilization hits 100%, latency spikes, and requests start timing out globally. How do you scale LLM inference infrastructure during sudden traffic spikes?


What you need to know

Why a queue with no limit makes things worse

When requests arrive faster than GPUs can serve them, an unbounded queue grows without end. Every request waits behind every other one, and soon every request times out, including ones that could have been served. A bounded queue rejects extra requests at once with 503 Service Unavailable and a Retry-After header. Some users get a fast "try again"; the rest get real answers.

Why GPU utilisation is the wrong signal

GPU utilisation reaches 100% under moderate load and stays there while latency keeps getting worse, so it cannot tell you how overloaded you are. Scale on what users feel: the number of requests waiting, and time to first token (TTFT). vLLM, for example, exports vllm:num_requests_waiting for Prometheus.

YAML
apiVersion: keda.sh/v1alpha1kind: ScaledObjectmetadata:  name: llm-serverspec:  scaleTargetRef:    name: vllm-deployment  minReplicaCount: 4  maxReplicaCount: 40  triggers:    - type: prometheus      metadata:        serverAddress: http://prometheus.monitoring:9090        query: sum(vllm:num_requests_waiting)        threshold: "8"

This KEDA rule adds replicas when more than about eight requests per replica are waiting. It only helps if new replicas start fast: pre-pull images onto a warm node pool and keep model weights on local NVMe, so a new replica serves in about a minute instead of ten.

The plan by time

WhenMoveWhy
First minutesPer-key and per-IP limits, bounded queue, priority tiersStops universal timeouts
First minutesBehind flags: lower max_tokens, turn off very long context and n greater than 1 samplingEach request costs less GPU time
First hourWarm pool, scale on queue depth and TTFTAdds capacity that arrives quickly
First hourSecond region, hosted API fallback for the same model familyAvailability beats margin during a spike
First daysPrefix caching, semantic cache, tuned batchingCuts cost per request

Efficiency once the fire is out

Serving engines such as vLLM, SGLang and TensorRT-LLM use continuous batching (new requests join the running batch between steps, instead of waiting for a whole batch to finish) and paged KV cache (memory for each request's attention cache is allocated in small pages, so more requests fit). Prefix caching reuses the processed system prompt across requests. A semantic cache answers repeated questions from stored answers; viral traffic is unusually repetitive, so hit rates are high.

Watch TTFT and inter-token latency p95, queue wait, rejection rate, tokens per second per GPU, and cost per 1,000 requests.

A real-life example

Scenario, numbers made up. An exam-prep app launches an "explain my answer" feature the night board results come out. Traffic rises 12 times in 40 minutes. The queue has no limit, so p95 latency passes 90 seconds and 70% of requests time out, including those from paying subscribers.

On-call caps the queue at 200 per replica, returns 503 with Retry-After: 15 beyond that, and gives subscribers a priority lane. They cut max_tokens from 1,200 to 600 behind a flag. Timeouts fall to 6% within 20 minutes. Meanwhile the warm pool brings replicas from 8 to 30, and a hosted API takes the overflow. Next morning, prefix caching and a semantic cache (hit rate 38%, since students ask the same questions) cut GPU need by about a third.

Follow-up questions to expect

  • "What does a semantic cache risk?" — Returning a stored answer to a question that only looks similar. Use a strict similarity threshold, key the cache by tenant and model version, and skip it for personal data.
  • "Why not just scale on CPU or GPU utilisation like a web app?" — It saturates before latency shows the damage, so it is too late and too coarse. Queue depth and TTFT track what users feel.
  • "What if the hosted fallback gives different answers?" — Keep a tested prompt for it and log which path served each answer. During a spike, a slightly different answer beats a timeout.