Building AI Features in Python Backends

Retries, backoff, fallbacks and circuit breakers


On a Tuesday evening, ShipFast's model provider returned "overloaded" (HTTP 529) for about 30% of requests for 14 minutes. The triage workers had a simple retry: on any error, try again immediately, up to five times. Each failed message made six calls in under a second. The workers' traffic to the provider tripled, which made the overload worse for everyone, including ShipFast. Then the provider started returning 429 rate-limit errors too. By the end of the incident, 2,300 messages had failed all five retries and were marked failed, and the day's bill showed $140 of extra calls.

Every piece of that story is a known pattern with a known fix. The model is just another remote dependency, and it fails the ways remote dependencies fail. This lesson adds the standard defences in layers, each doing one job.

Four layers, each catching what the one above cannotRetry: backoff, jitter, 20 s deadlineBreaker: 5 fails in 30 s, 20 s pauseFallback model, same complete()Degrade: unknown, human triage, flagged
Retries fix blips, the breaker stops hammering a provider that is down, and nothing in the stack ever retries a bad request.

The layers

  1. Retry with backoff — wait longer after each failure, with random jitter, and stop at a total deadline.
  2. Circuit breaker — after several failures in a short window, stop calling the primary for a cooldown period.
  3. Fallback model — send the call to a second model while the primary is failing or its circuit is open.
  4. Degraded result — if all else fails, return a safe answer: unknown, human triage, degraded: true.

Each layer handles what the one before it cannot. Retries fix brief blips. The breaker stops retries from hammering a provider that is really down. The fallback keeps work flowing during a longer outage. Degradation makes sure a message is never lost or stuck even when both models fail.

Retries that help instead of hurt

Python
# shipfast/resilience.pyimport asyncioimport randomimport timefrom typing import Awaitable, Callablefrom shipfast.llm import LLMClient, LLMResult, LLMTimeout, LLMUnavailableRETRYABLE = (LLMTimeout, LLMUnavailable)async def with_retries(call: Callable[[], Awaitable[LLMResult]], *, attempts: int = 3,                       base_s: float = 0.5, cap_s: float = 4.0,                       deadline_s: float = 20.0) -> LLMResult:    start = time.monotonic()    for attempt in range(1, attempts + 1):        try:            return await call()        except RETRYABLE:            delay = random.uniform(0, min(cap_s, base_s * 2 ** attempt))   # full jitter            out_of_time = time.monotonic() - start + delay > deadline_s            if attempt == attempts or out_of_time:                raise            await asyncio.sleep(delay)    raise AssertionError("unreachable")

This is where the error mapping from Section 2 pays off. Only LLMTimeout and LLMUnavailable are retried. LLMBadRequest passes straight through, because a malformed request fails the same way every time.

Exponential backoff doubles the maximum wait after each failure: up to 1 second, then 2, capped at 4. Full jitter picks a random wait between zero and that maximum, so a thousand workers that failed at the same moment do not all retry at the same moment. The deadline caps total time, so a background job cannot spend minutes on one message. Three attempts is enough: if the provider fails three times across a few seconds, the problem is not a blip.

The circuit breaker

Retries are per call. When the provider is down for minutes, every call still tries three times before giving up, and all those attempts add load to a provider that is already struggling. A circuit breaker watches failures across all calls and, past a threshold, stops trying for a while.

Python
# shipfast/resilience.py (continued)class CircuitBreaker:    def __init__(self, threshold: int = 5, window_s: float = 30.0, cooldown_s: float = 20.0):        self.threshold, self.window_s, self.cooldown_s = threshold, window_s, cooldown_s        self.failures: list[float] = []        self.opened_at: float | None = None    def allow(self) -> bool:        if self.opened_at is None:            return True        return time.monotonic() - self.opened_at >= self.cooldown_s   # half-open: try again    def record_success(self) -> None:        self.failures.clear()        self.opened_at = None    def record_failure(self) -> None:        now = time.monotonic()        self.failures = [t for t in self.failures if now - t < self.window_s] + [now]        if len(self.failures) >= self.threshold:            self.opened_at = now

Five failed calls within 30 seconds open the circuit. For the next 20 seconds, allow() returns false and calls go straight to the fallback without touching the primary. After the cooldown, calls are allowed through again (the "half-open" state). One success closes the circuit; another failure, with the recent failures still in the window, opens it again at once. This is a small, per-process breaker. Across many instances, each keeps its own, which is usually fine; a shared breaker in Redis is possible but rarely worth the complexity.

The fallback, behind the same interface

ResilientLLM combines the pieces and offers the same complete() as LLMClient, so nothing that calls the model needs to change.

Python
# shipfast/resilience.py (continued)class ResilientLLM:    """Same complete() as LLMClient: retries, then a breaker, then a fallback model."""    def __init__(self, primary: LLMClient, fallback: LLMClient | None = None,                 breaker: CircuitBreaker | None = None) -> None:        self.primary, self.fallback = primary, fallback        self.breaker = breaker or CircuitBreaker()        self.model = primary.model    async def complete(self, **kwargs) -> LLMResult:        if self.breaker.allow():            try:                result = await with_retries(lambda: self.primary.complete(**kwargs))                self.breaker.record_success()                return result            except RETRYABLE:                self.breaker.record_failure()                if self.fallback is None:                    raise        if self.fallback is None:            raise LLMUnavailable("primary circuit open, no fallback")        return await with_retries(lambda: self.fallback.complete(**kwargs), attempts=2)

This is where the LLM Protocol at the bottom of shipfast/llm.py comes in. It says "anything with a model attribute and this complete() signature". call_structured, classify and extract are typed against LLM, so they accept an LLMClient, a ResilientLLM or the test fake without a change.

Python
# shipfast/llm.py (end of file)class LLM(Protocol):    model: str    async def complete(self, *, feature: str, system: str, messages: list[dict],                       max_tokens: int = 512, schema: dict | None = None) -> LLMResult: ...

In shipfast/api.py, get_llm() now builds the resilient client. The streaming endpoint from Section 2 keeps using the primary directly, because a half-sent stream cannot be retried or moved to another model.

Python
@lru_cachedef get_llm() -> ResilientLLM:    primary = LLMClient(settings.llm_model, timeout_s=settings.llm_timeout_s)    fallback = LLMClient(settings.llm_fallback_model, timeout_s=settings.llm_timeout_s)    return ResilientLLM(primary, fallback)def get_stream_llm() -> LLMClient:    return get_llm().primary          # a half-sent stream cannot be retried or failed over

The streaming endpoint's dependency changes from get_llm to get_stream_llm; nothing else in it changes.

When everything fails

If both models fail, complete() raises LLMUnavailable. The triage pipeline catches it and returns a degraded result: intent unknown, queue human_triage, degraded: true. The message is never lost and never stuck, and agents handle it the way they did before the service existed. When classification succeeded but a later call failed, the pipeline keeps the label and routes on it, with no draft. The next lesson's triage() shows both paths.

Check your understanding

0 of 3 answered

1.Why does with_retries use a random delay instead of a fixed doubling delay?

2.The primary model's circuit is open. What does ResilientLLM.complete() do with a new call?

3.Why does the streaming draft endpoint not use ResilientLLM?