Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

How do you engineer a production integration with an LLM provider's API: timeouts, retries, rate limits and fallbacks?


What you need to know

An LLM API is a remote service that is slow, rate-limited, sometimes overloaded, and billed per token. The engineering around the call decides whether a provider incident becomes a slower answer or an outage.

What the internal client does

ConcernRuleWhy
TimeoutsSeparate limits for streaming and non-streaming callsA hung connection must not hold a worker for minutes
RetriesOnly on 429, 5xx (including "overloaded") and connection errorsThese are temporary; a 400 will fail the same way again
BackoffExponential with jitter, with a total time budgetJitter stops all clients retrying at the same moment
Rate limitsRead remaining-quota and retry-after headersSlow down before hitting 429, not after
IdempotencyKeys on any downstream action the reply triggersA retry after a timeout must not send an email twice
FallbackCircuit breaker to a second model or a cached/degraded replyUsers get a worse answer, not an error page
ObservabilityLog model, prompt version, tokens, latency, cost, trace idCost becomes a metric, not a monthly surprise

Retries: configure, don't double-wrap

The official OpenAI and Anthropic Python SDKs already retry some errors, twice by default, with backoff. If you add your own retry loop around them, three attempts can quietly become nine. Set the SDK's max_retries and timeout deliberately, and keep one retry policy.

Python
import timefrom anthropic import Anthropic, APIStatusErrorclient = Anthropic(max_retries=3, timeout=30.0)          # one retry policy, set explicitlydef complete(messages, *, prompt_version: str, **kw):    start = time.monotonic()    try:        resp = client.messages.create(model=PRIMARY, max_tokens=1024, messages=messages, **kw)    except APIStatusError as e:        if e.status_code >= 500 and breaker.record_failure():            return fallback_complete(messages, **kw)     # second provider or degraded answer        raise    log_call(model=PRIMARY, prompt_version=prompt_version, usage=resp.usage,             latency_ms=int((time.monotonic() - start) * 1000))    return resp

The client object carries the retry and timeout policy. The wrapper adds what the SDK cannot know: the breaker, the fallback and the cost log.

Streaming details

Stream user-facing calls, because time to first token is what users feel. Handle a disconnect in the middle of a stream: save the partial output, and do not retry automatically in a way that shows the user a second, different answer.

Rolling it out

  1. Day one — the wrapper with timeouts, one retry policy and token and cost logging.
  2. Week one — dashboards for error rate by status code, p95 latency and cost per feature.
  3. After a week of data — set breaker thresholds and wire in a tested fallback model.
  4. Ongoing — per-feature budgets and alerts on spend and error spikes.

A real-life example

Scenario (illustrative numbers). A ride-hailing company's driver-support bot calls a model API directly from six services, each with its own retry code. During a 40-minute provider incident, some services retry five times with no backoff; request volume to the provider rises 4x, and the bot is down for the whole incident.

The team replaces the six call sites with one internal client: 20-second timeout, SDK retries set to 2 with jitter, a breaker that opens at 30% errors over one minute, and a fallback to a second provider's model already tested on the eval set. In the next provider incident, the breaker opens within a minute, 96% of conversations complete on the fallback, and the cost log shows the incident added about ₹4,000 in spend.

Follow-up questions to expect

  • "How do you choose the fallback model?" — Run it on the same eval set in advance, keep its prompts maintained, and send it a small share of traffic regularly so you know it still works.
  • "What do you do about a 429 that keeps coming?" — Respect retry-after, queue non-urgent work, and ask for a higher limit or spread load across regions or providers.
  • "Should you cache responses?" — For repeated identical requests, yes; also use provider prompt caching for long static prefixes.