Course Content
LLMOps & Deployment
6 sections · 40 lessons
How do you design systems to handle provider outages gracefully?
What you need to know
Python
1import time23class CircuitBreaker:4 def __init__(self, max_failures=5, cooldown=30.0):5 self.max_failures, self.cooldown = max_failures, cooldown6 self.failures, self.opened_at = 0, None78 def state(self) -> str:9 if self.opened_at is None:10 return "closed"11 if time.monotonic() - self.opened_at >= self.cooldown:12 return "half_open" # let one probe request through13 return "open"1415 def allow(self) -> bool:16 return self.state() != "open"1718 def record(self, ok: bool) -> None:19 if ok:20 self.failures, self.opened_at = 0, None # close again21 else:22 self.failures += 123 if self.state() == "half_open" or self.failures >= self.max_failures:24 self.opened_at = time.monotonic() # (re)openBefore each call, the gateway checks allow(); if false, it goes straight to the next provider. This simple version lets several probes through in half-open state; production breakers (in gateways or libraries) limit it to one and usually use error rates over a time window rather than a raw count.
The rest of the design
- Real health checks. A small completion every 30 s per provider and region, measuring latency and output, not just an HTTP ping.
- Pre-provisioned failover. Keys, quota and evaluated prompt variants for the second route, already tested with live traffic. You cannot sign a contract during an outage.
- Idempotency keys. If a request triggers an action (sending an SMS, raising a refund ticket), retries must not repeat it.
- Queue asynchronous work. Summaries, emails and batch jobs wait and replay; for them an outage is only a delay.
- Honest UI. "Our assistant is running in limited mode" is better than a spinner or an invented answer.
- Game days. Deliberately block the primary provider in staging, and sometimes in production for a small slice, to prove the path works.
A real-life example
On the second night of a Diwali sale, a fintech support bot's primary provider starts timing out on 30% of calls in one region at 9:12 pm.
- The circuit breaker for that region opens after 5 failures in 10 seconds; traffic moves to the provider's second region, which has reserved quota.
- At 9:20 pm the second region also degrades. That breaker opens, and the gateway fails over to the same model family on a second cloud platform — a path that has carried 2% of traffic every day for three months, so its prompt and parser are known to work.
- Refund-ticket creation uses idempotency keys, so the 400 requests that were retried during the switch do not create duplicate tickets.
- Order-status questions use the deterministic API path and are unaffected.
Users see slightly slower answers for about eight minutes. The dashboard records 14% of traffic on the failover path for two hours, and the provider's incident report arrives the next morning.
Follow-up questions to expect
- "How do you set the breaker thresholds?" — From normal error rates: open well above the everyday level (for example over 20% errors in 30 seconds with a minimum request count), and cool down for 30–60 seconds.
- "What if all providers are down?" — The degradation ladder: cached answers, deterministic paths, retrieval-only answers, then an honest failure message.
- "Why idempotency keys for LLM calls?" — The LLM call itself is harmless to repeat, but the actions the agent takes after it — tickets, messages, payments — are not.