Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
How do you engineer a production integration with an LLM provider's API: timeouts, retries, rate limits and fallbacks?
What you need to know
An LLM API is a remote service that is slow, rate-limited, sometimes overloaded, and billed per token. The engineering around the call decides whether a provider incident becomes a slower answer or an outage.
What the internal client does
| Concern | Rule | Why |
|---|---|---|
| Timeouts | Separate limits for streaming and non-streaming calls | A hung connection must not hold a worker for minutes |
| Retries | Only on 429, 5xx (including "overloaded") and connection errors | These are temporary; a 400 will fail the same way again |
| Backoff | Exponential with jitter, with a total time budget | Jitter stops all clients retrying at the same moment |
| Rate limits | Read remaining-quota and retry-after headers | Slow down before hitting 429, not after |
| Idempotency | Keys on any downstream action the reply triggers | A retry after a timeout must not send an email twice |
| Fallback | Circuit breaker to a second model or a cached/degraded reply | Users get a worse answer, not an error page |
| Observability | Log model, prompt version, tokens, latency, cost, trace id | Cost becomes a metric, not a monthly surprise |
Retries: configure, don't double-wrap
The official OpenAI and Anthropic Python SDKs already retry some errors, twice by default, with backoff. If you add your own retry loop around them, three attempts can quietly become nine. Set the SDK's max_retries and timeout deliberately, and keep one retry policy.
1import time2from anthropic import Anthropic, APIStatusError34client = Anthropic(max_retries=3, timeout=30.0) # one retry policy, set explicitly56def complete(messages, *, prompt_version: str, **kw):7 start = time.monotonic()8 try:9 resp = client.messages.create(model=PRIMARY, max_tokens=1024, messages=messages, **kw)10 except APIStatusError as e:11 if e.status_code >= 500 and breaker.record_failure():12 return fallback_complete(messages, **kw) # second provider or degraded answer13 raise14 log_call(model=PRIMARY, prompt_version=prompt_version, usage=resp.usage,15 latency_ms=int((time.monotonic() - start) * 1000))16 return respThe client object carries the retry and timeout policy. The wrapper adds what the SDK cannot know: the breaker, the fallback and the cost log.
Streaming details
Stream user-facing calls, because time to first token is what users feel. Handle a disconnect in the middle of a stream: save the partial output, and do not retry automatically in a way that shows the user a second, different answer.
Rolling it out
- Day one — the wrapper with timeouts, one retry policy and token and cost logging.
- Week one — dashboards for error rate by status code, p95 latency and cost per feature.
- After a week of data — set breaker thresholds and wire in a tested fallback model.
- Ongoing — per-feature budgets and alerts on spend and error spikes.
A real-life example
Scenario (illustrative numbers). A ride-hailing company's driver-support bot calls a model API directly from six services, each with its own retry code. During a 40-minute provider incident, some services retry five times with no backoff; request volume to the provider rises 4x, and the bot is down for the whole incident.
The team replaces the six call sites with one internal client: 20-second timeout, SDK retries set to 2 with jitter, a breaker that opens at 30% errors over one minute, and a fallback to a second provider's model already tested on the eval set. In the next provider incident, the breaker opens within a minute, 96% of conversations complete on the fallback, and the cost log shows the incident added about ₹4,000 in spend.
Follow-up questions to expect
- "How do you choose the fallback model?" — Run it on the same eval set in advance, keep its prompts maintained, and send it a small share of traffic regularly so you know it still works.
- "What do you do about a 429 that keeps coming?" — Respect
retry-after, queue non-urgent work, and ask for a higher limit or spread load across regions or providers. - "Should you cache responses?" — For repeated identical requests, yes; also use provider prompt caching for long static prefixes.