AI Product Engineering: Shipping LLM Features That Last

Course Content

AI Product Engineering: Shipping LLM Features That Last

6 sections · 22 lessons

Surviving outages: timeouts, fallbacks, degraded modes


On the Friday before Diwali, at 8:15 pm, the model provider's error rate climbed. About 30% of calls returned "overloaded" errors, and many of the rest took 20 seconds or more. TiffinGo's ticket volume was already double the normal rate. Drafts stopped appearing, and 45 agents stared at spinning placeholders while the queue passed 2,000 tickets.

The assistant had been built on the assumption that the model is always there. It is not. Every hosted API has incidents, and they tend to come at peak times, because peak times are when everyone else is busy too.

This lesson is about the plan for when the model is slow, wrong or gone. The goal is not to never fail. It is that when the model fails, agents lose some speed, not their ability to work.

The fallback chain on a bad nightPrimary release,behind a breakerFallback release,already gate-approvedDegraded mode:code-only screenRedraft whenthe providerrecoversFive failures in 30 seconds open the breaker for 60.
Most of the pipeline was never the model, so an outage costs agents speed — 3.9 minutes a ticket instead of 2.3 — not the ability to work.

How a model dependency fails

FailureWhat you seeHow to detect it
Hard errors5xx or "overloaded" responsesError rate per minute, by error type
Rate limiting429 responsesCount of 429s; compare with your quota
Slownessp95 latency from 5 s to 30 s or moreLatency percentiles, not averages
Truncation or refusalstop_reason is max_tokens or refusalCounted in the client, never ignored
Quiet quality changeValid drafts that are more often wrongApproval and validator trends (next lesson)

The first four are loud and fast. The last is quiet and slow, and needs monitoring rather than error handling. This lesson handles the loud ones.

Timeouts and retries that do not make things worse

A timeout protects your service from waiting forever. A retry recovers from brief failures. Together, set badly, they turn a provider incident into your incident: every stuck call holds a worker, and every retry adds load to a provider that is already overloaded.

TiffinGo's rules:

  • Per-call timeout: 10 seconds. The p95 is 5.1 seconds, so 10 seconds cuts off only calls that are clearly stuck.
  • One retry, only for retryable errors. Network errors, 429s and 5xx responses. Never retry a 400; it is a bug and will fail again.
  • Total budget per draft: about 30 seconds, including the repair and the fallback. With a 10-second timeout, that allows the primary and one fallback call, not a long chain. Drafts are generated when the ticket arrives, so this is still far below the queue wait.
  • Respect rate-limit hints. The SDK honours the provider's retry-after header on 429s, which is why the client uses the SDK's own single retry rather than a loop of its own.

The client from section 3 also needs to tell the rest of the service when the provider is the problem. It gains one exception type:

Python
# additions to llm.pyclass Unavailable(LLMError):    """The provider could not serve the call: network, timeout, rate limit or server error."""RETRYABLE = (anthropic.APIConnectionError, anthropic.RateLimitError, anthropic.InternalServerError)# inside complete(), around the messages.create call:#     try:#         resp = _client.messages.create(...)#     except RETRYABLE as exc:#         raise Unavailable(type(exc).__name__) from exc

APIConnectionError covers network failures and timeouts, RateLimitError covers 429, and InternalServerError covers 500 and above, including "overloaded". Only llm.py knows these provider classes; the rest of the service sees llm.Unavailable.

A circuit breaker

When the provider is failing, calling it for every ticket wastes 10 seconds per ticket and adds to the overload. A circuit breaker notices repeated failures and stops calling for a while.

Python
import timeclass CircuitBreaker:    def __init__(self, failure_limit: int = 5, window_s: float = 30, cooldown_s: float = 60):        self.failure_limit, self.window_s, self.cooldown_s = failure_limit, window_s, cooldown_s        self.failures: list[float] = []        self.opened_at: float | None = None    def allow(self) -> bool:        if self.opened_at is None:            return True        if time.monotonic() - self.opened_at >= self.cooldown_s:            self.opened_at, self.failures = None, []  # try the provider again            return True        return False    def record(self, ok: bool) -> None:        now = time.monotonic()        if ok:            self.failures = []            return        self.failures = [t for t in self.failures if now - t < self.window_s] + [now]        if len(self.failures) >= self.failure_limit:            self.opened_at = now

Five failures within 30 seconds open the breaker. For the next 60 seconds, calls skip that model entirely. After the cooldown, calls are allowed again; if they fail again, it reopens within seconds. In a multi-worker service, keep this state in a shared store such as Redis, or each worker learns about the outage separately.

The fallback chain

Python
PRIMARY = load_release("v6")            # medium tier, the normal releaseFALLBACK = load_release("v6-fallback")  # same contract on a different, gate-approved modelBREAKERS = {PRIMARY.version: CircuitBreaker(), FALLBACK.version: CircuitBreaker()}def draft_with_fallback(ticket_text: str, order: dict, minutes_late: int):    for release in (PRIMARY, FALLBACK):        breaker = BREAKERS[release.version]        if not breaker.allow():            continue        try:            draft = draft_for_ticket(release, ticket_text, order, minutes_late)            breaker.record(ok=True)            return draft, release.version        except llm.Unavailable:            breaker.record(ok=False)   # provider problem: try the next release        except llm.LLMError:            return None, release.version  # refusal or truncation: manual path, not an outage    return None, "degraded"

The fallback is a full prompt release, not a model name swapped at runtime. Prompts often behave differently on a different model, so v6-fallback has its own folder, passed the regression gate, and is re-tested whenever v6 changes. TiffinGo's fallback uses the large-tier model, which scored 116 of 120. It costs more, but only during incidents.

A fallback on the same provider helps when one model is overloaded, which is the most common incident. It does not help when the whole provider is down. A second provider covers that, at the cost of a second adapter in llm.py, a second prompt release to maintain, and a second data-processing agreement. TiffinGo decided the degraded mode below was enough for full outages, and wrote down that decision.

Degraded mode: useful without the model

When both releases fail, the ticket still opens, and much of the pipeline still works, because most of it was never the model. The late-delivery check (R4), the keyword safety net, the R6 and R7 limits, and policy retrieval are all ordinary code.

So the degraded screen shows: a banner saying "Assistant unavailable, drafts will return", the order with paid prices, any R4 refund already computed, any forced escalation, and the three most relevant policy sections for this city. The agent works the ticket by hand, but faster than before the assistant existed. The ticket is also queued, so if it is still open when the provider recovers, a draft appears.

Check your understanding

0 of 3 answered

1.During a provider overload, why is "retry every failed call up to five times" a bad policy?

2.Why is the fallback a separate prompt release instead of the same prompt with the model name changed at runtime?

3.Both the primary and fallback releases are failing. What should the agent see?