Live Coding Interview Prep

Course Content

Live Coding Interview Prep

7 sections · 50 lessons

Implement fallback between multiple LLM providers.


The primary provider's circuit breakerclosed:calls go throughopen:skipped instantlyhalf-open:one probeafter 2 timeoutsafter 60 scooldownAt 2 s the request went straight to the backup; at 61 s one probe succeeded and closed the circuit.
The breaker is what stops every request paying a dead provider's full timeout before failing over.

What you need to know

Providers have outages and rate-limit spikes. Fallback keeps your product up by trying another model. Two ideas make it work well:

Classify the failure first. Fail over only when another provider could succeed:

errorfail over?why
timeout, connection error, 429, 500–529yesthe problem is that provider
400 bad request, prompt too longnothe request itself is wrong; it fails everywhere
401 / 403no, alertyour credentials or account

A circuit breaker remembers failures. It has three states:

  1. Closed — requests go through; count consecutive failures.
  2. Open — after N failures, skip this provider completely until a cooldown passes.
  3. Half-open — after the cooldown, let exactly one probe request through. Success closes the circuit; failure opens it again.

Without a breaker, a provider that times out after 30 seconds costs every request 30 seconds before failing over. With it, only the first few requests pay.

Separately from outages, the Anthropic API has a server-side fallbacks option for a different case: when a model declines a request (stop_reason: "refusal"), it can reroute to another model. It does not cover outages, so you still need this chain.

Python
import timefrom collections.abc import Callableclass CircuitBreaker:    """closed -> open after `threshold` failures -> half-open after `cooldown`."""    def __init__(self, threshold: int = 5, cooldown: float = 30.0,                 clock: Callable[[], float] = time.monotonic) -> None:        self.threshold, self.cooldown, self.clock = threshold, cooldown, clock        self.failures, self.opened_at, self.probing = 0, None, False    def allow(self) -> bool:        if self.opened_at is None:            return True        if not self.probing and self.clock() - self.opened_at >= self.cooldown:            self.probing = True                  # half-open: exactly one probe            return True        return False    def record(self, ok: bool) -> None:        if ok:            self.failures, self.opened_at, self.probing = 0, None, False        else:            self.failures += 1            if self.probing or self.failures >= self.threshold:                self.opened_at, self.probing = self.clock(), Falsedef is_retryable(exc: Exception) -> bool:    status = getattr(exc, "status_code", None)    return isinstance(exc, (TimeoutError, ConnectionError)) or status in {408, 429} or (        status is not None and status >= 500)class ProviderChain:    """Try providers in order; skip ones whose circuit is open."""    def __init__(self, providers: list[tuple[str, Callable[[str], str]]], **breaker_kw) -> None:        self.providers = providers        self.breakers = {name: CircuitBreaker(**breaker_kw) for name, _ in providers}    def complete(self, prompt: str) -> dict:        errors = []        for name, call in self.providers:            breaker = self.breakers[name]            if not breaker.allow():                errors.append(f"{name}: circuit open")                continue            try:                text = call(prompt)            except Exception as exc:                if not is_retryable(exc):                    raise                        # a bad request fails on every provider                breaker.record(False)                errors.append(f"{name}: {type(exc).__name__}")                continue            breaker.record(True)            return {"provider": name, "text": text}        raise RuntimeError("all providers failed: " + "; ".join(errors))

The tricky parts:

  • probing is what makes half-open mean one probe. Without it, the moment the cooldown passes every waiting request rushes to the provider that may still be down.
  • A failed probe re-opens immediately, even though the failure count was not reset; a half-open circuit gets one chance.
  • Non-retryable errors are re-raised without touching the breaker. A burst of malformed requests should not open the circuit on a healthy provider.
  • The clock is injectable so the trace below can move time forward.

Complexity: at most p provider calls per request, and each breaker check is O(1). With an open circuit, a down provider costs one dictionary lookup instead of a timeout.

A real-life example

Python
now = [0.0]state = {"primary_up": False}def primary(prompt):    if not state["primary_up"]:        raise TimeoutError("30s timeout")    return "primary answer"backup = lambda prompt: "backup answer"chain = ProviderChain([("primary", primary), ("backup", backup)],                      threshold=2, cooldown=60, clock=lambda: now[0])for t in [0, 1, 2, 61, 62]:    now[0] = float(t)    if t == 61:        state["primary_up"] = True    print(t, chain.complete("hi")["provider"], chain.breakers["primary"].failures)# 0 backup 1# 1 backup 2# 2 backup 2# 61 primary 0# 62 primary 0
timeprimary breaker beforeprimary called?served by
0 sclosed, 0 failuresyes, times out → 1backup
1 sclosed, 1 failureyes, times out → 2, opensbackup
2 sopenno — skipped instantlybackup
61 scooldown passed → half-openyes, one probe, succeeds → closedprimary
62 sclosedyesprimary

At 2 s the user did not wait for a timeout at all. That is the difference the breaker makes during a real outage.

A fintech chatbot that must answer "why was my UPI payment declined?" during a provider outage uses a chain like this, with a smaller model from a second provider as the backup.

Follow-up questions to expect

  • "Is the backup's answer as good?" — Usually not; prompts are tuned for the primary. Run your eval suite against the backup too, and log which provider served each response.
  • "What about streaming?" — If the primary fails after sending some tokens, the user sees text restart. Fail over only before the first token, or tell the client to clear and restart.
  • "How do breakers work with many server processes?" — Each process has its own; that is usually fine. For a shared view, keep breaker state in Redis with a short TTL.