Course Content
Live Coding Interview Prep
7 sections · 50 lessons
Implement fallback between multiple LLM providers.
What you need to know
Providers have outages and rate-limit spikes. Fallback keeps your product up by trying another model. Two ideas make it work well:
Classify the failure first. Fail over only when another provider could succeed:
| error | fail over? | why |
|---|---|---|
| timeout, connection error, 429, 500–529 | yes | the problem is that provider |
| 400 bad request, prompt too long | no | the request itself is wrong; it fails everywhere |
| 401 / 403 | no, alert | your credentials or account |
A circuit breaker remembers failures. It has three states:
- Closed — requests go through; count consecutive failures.
- Open — after N failures, skip this provider completely until a cooldown passes.
- Half-open — after the cooldown, let exactly one probe request through. Success closes the circuit; failure opens it again.
Without a breaker, a provider that times out after 30 seconds costs every request 30 seconds before failing over. With it, only the first few requests pay.
Separately from outages, the Anthropic API has a server-side fallbacks option for a different case: when a model declines a request (stop_reason: "refusal"), it can reroute to another model. It does not cover outages, so you still need this chain.
1import time2from collections.abc import Callable34class CircuitBreaker:5 """closed -> open after `threshold` failures -> half-open after `cooldown`."""67 def __init__(self, threshold: int = 5, cooldown: float = 30.0,8 clock: Callable[[], float] = time.monotonic) -> None:9 self.threshold, self.cooldown, self.clock = threshold, cooldown, clock10 self.failures, self.opened_at, self.probing = 0, None, False1112 def allow(self) -> bool:13 if self.opened_at is None:14 return True15 if not self.probing and self.clock() - self.opened_at >= self.cooldown:16 self.probing = True # half-open: exactly one probe17 return True18 return False1920 def record(self, ok: bool) -> None:21 if ok:22 self.failures, self.opened_at, self.probing = 0, None, False23 else:24 self.failures += 125 if self.probing or self.failures >= self.threshold:26 self.opened_at, self.probing = self.clock(), False2728def is_retryable(exc: Exception) -> bool:29 status = getattr(exc, "status_code", None)30 return isinstance(exc, (TimeoutError, ConnectionError)) or status in {408, 429} or (31 status is not None and status >= 500)3233class ProviderChain:34 """Try providers in order; skip ones whose circuit is open."""3536 def __init__(self, providers: list[tuple[str, Callable[[str], str]]], **breaker_kw) -> None:37 self.providers = providers38 self.breakers = {name: CircuitBreaker(**breaker_kw) for name, _ in providers}3940 def complete(self, prompt: str) -> dict:41 errors = []42 for name, call in self.providers:43 breaker = self.breakers[name]44 if not breaker.allow():45 errors.append(f"{name}: circuit open")46 continue47 try:48 text = call(prompt)49 except Exception as exc:50 if not is_retryable(exc):51 raise # a bad request fails on every provider52 breaker.record(False)53 errors.append(f"{name}: {type(exc).__name__}")54 continue55 breaker.record(True)56 return {"provider": name, "text": text}57 raise RuntimeError("all providers failed: " + "; ".join(errors))The tricky parts:
probingis what makes half-open mean one probe. Without it, the moment the cooldown passes every waiting request rushes to the provider that may still be down.- A failed probe re-opens immediately, even though the failure count was not reset; a half-open circuit gets one chance.
- Non-retryable errors are re-raised without touching the breaker. A burst of malformed requests should not open the circuit on a healthy provider.
- The clock is injectable so the trace below can move time forward.
Complexity: at most p provider calls per request, and each breaker check is O(1). With an open circuit, a down provider costs one dictionary lookup instead of a timeout.
A real-life example
1now = [0.0]2state = {"primary_up": False}3def primary(prompt):4 if not state["primary_up"]:5 raise TimeoutError("30s timeout")6 return "primary answer"7backup = lambda prompt: "backup answer"89chain = ProviderChain([("primary", primary), ("backup", backup)],10 threshold=2, cooldown=60, clock=lambda: now[0])11for t in [0, 1, 2, 61, 62]:12 now[0] = float(t)13 if t == 61:14 state["primary_up"] = True15 print(t, chain.complete("hi")["provider"], chain.breakers["primary"].failures)16# 0 backup 117# 1 backup 218# 2 backup 219# 61 primary 020# 62 primary 0| time | primary breaker before | primary called? | served by |
|---|---|---|---|
| 0 s | closed, 0 failures | yes, times out → 1 | backup |
| 1 s | closed, 1 failure | yes, times out → 2, opens | backup |
| 2 s | open | no — skipped instantly | backup |
| 61 s | cooldown passed → half-open | yes, one probe, succeeds → closed | primary |
| 62 s | closed | yes | primary |
At 2 s the user did not wait for a timeout at all. That is the difference the breaker makes during a real outage.
A fintech chatbot that must answer "why was my UPI payment declined?" during a provider outage uses a chain like this, with a smaller model from a second provider as the backup.
Follow-up questions to expect
- "Is the backup's answer as good?" — Usually not; prompts are tuned for the primary. Run your eval suite against the backup too, and log which provider served each response.
- "What about streaming?" — If the primary fails after sending some tokens, the user sees text restart. Fail over only before the first token, or tell the client to clear and restart.
- "How do breakers work with many server processes?" — Each process has its own; that is usually fine. For a shared view, keep breaker state in Redis with a short TTL.