LLMOps & Deployment

Course Content

LLMOps & Deployment

6 sections · 40 lessons

How do you design systems to handle provider outages gracefully?


Circuit breaker states for one provider regionclosed:calls passopen: fail fasthalf-open:one probeafter 5failuresafter cooldownA successful probe closes the circuit; a failed probe opens it again.
Failing fast in the open state is what stops a partial provider outage from spending every request's timeout budget.

What you need to know

Python
import timeclass CircuitBreaker:    def __init__(self, max_failures=5, cooldown=30.0):        self.max_failures, self.cooldown = max_failures, cooldown        self.failures, self.opened_at = 0, None    def state(self) -> str:        if self.opened_at is None:            return "closed"        if time.monotonic() - self.opened_at >= self.cooldown:            return "half_open"            # let one probe request through        return "open"    def allow(self) -> bool:        return self.state() != "open"    def record(self, ok: bool) -> None:        if ok:            self.failures, self.opened_at = 0, None        # close again        else:            self.failures += 1            if self.state() == "half_open" or self.failures >= self.max_failures:                self.opened_at = time.monotonic()           # (re)open

Before each call, the gateway checks allow(); if false, it goes straight to the next provider. This simple version lets several probes through in half-open state; production breakers (in gateways or libraries) limit it to one and usually use error rates over a time window rather than a raw count.

The rest of the design

  • Real health checks. A small completion every 30 s per provider and region, measuring latency and output, not just an HTTP ping.
  • Pre-provisioned failover. Keys, quota and evaluated prompt variants for the second route, already tested with live traffic. You cannot sign a contract during an outage.
  • Idempotency keys. If a request triggers an action (sending an SMS, raising a refund ticket), retries must not repeat it.
  • Queue asynchronous work. Summaries, emails and batch jobs wait and replay; for them an outage is only a delay.
  • Honest UI. "Our assistant is running in limited mode" is better than a spinner or an invented answer.
  • Game days. Deliberately block the primary provider in staging, and sometimes in production for a small slice, to prove the path works.

A real-life example

On the second night of a Diwali sale, a fintech support bot's primary provider starts timing out on 30% of calls in one region at 9:12 pm.

  • The circuit breaker for that region opens after 5 failures in 10 seconds; traffic moves to the provider's second region, which has reserved quota.
  • At 9:20 pm the second region also degrades. That breaker opens, and the gateway fails over to the same model family on a second cloud platform — a path that has carried 2% of traffic every day for three months, so its prompt and parser are known to work.
  • Refund-ticket creation uses idempotency keys, so the 400 requests that were retried during the switch do not create duplicate tickets.
  • Order-status questions use the deterministic API path and are unaffected.

Users see slightly slower answers for about eight minutes. The dashboard records 14% of traffic on the failover path for two hours, and the provider's incident report arrives the next morning.

Follow-up questions to expect

  • "How do you set the breaker thresholds?" — From normal error rates: open well above the everyday level (for example over 20% errors in 30 seconds with a minimum request count), and cool down for 30–60 seconds.
  • "What if all providers are down?" — The degradation ladder: cached answers, deterministic paths, retrieval-only answers, then an honest failure message.
  • "Why idempotency keys for LLM calls?" — The LLM call itself is harmless to repeat, but the actions the agent takes after it — tickets, messages, payments — are not.