LLMOps & Deployment

Course Content

LLMOps & Deployment

6 sections · 40 lessons

What fallback strategies can be used when a model fails or is unavailable?


What you need to know

The ladder

  1. Retry — only for transient errors, with exponential backoff and full jitter, honouring Retry-After; at most 2–3 attempts.
  2. Same model, other deployment — another region or account quota.
  3. Same model family, other provider — for example through a cloud marketplace such as Bedrock or Vertex.
  4. Smaller or older model — lower quality, still useful; prompt tuned and evaluated for it.
  5. Cached or deterministic answer — last known good answer, FAQ template, or a non-LLM path such as a status lookup.
  6. Honest failure — a clear message and a retry option, never an invented answer.

Retries done right

Exponential backoff doubles the wait after each failure (0.5 s, 1 s, 2 s…). Full jitter picks a random wait between zero and that ceiling, so thousands of clients do not retry at the same instant — a "thundering herd".

Python
import random, timeclass RetryableError(Exception):    def __init__(self, retry_after: float | None = None):        self.retry_after = retry_afterdef call_with_retries(fn, max_attempts=4, base=0.5, cap=8.0, budget=20.0):    start = time.monotonic()    for attempt in range(max_attempts):        try:            return fn()        except RetryableError as e:            if attempt == max_attempts - 1:                raise            # full jitter: random wait between 0 and the exponential ceiling            delay = random.uniform(0, min(cap, base * 2 ** attempt))            if e.retry_after is not None:                delay = max(delay, e.retry_after)   # respect the server            if time.monotonic() - start + delay > budget:                raise                               # out of time budget            time.sleep(delay)

Only retryable errors are retried: a 400 for a bad request will fail the same way every time. The budget keeps all retries inside the request's overall timeout, so the next tier of the ladder still has time.

What makes the ladder work

  • Circuit breaker per provider — after repeated failures, stop sending for a cooldown instead of paying a timeout on every request.
  • Evaluated fallback prompts — a prompt tuned for one model family often behaves differently on another. Each tier has its own prompt variant with eval results.
  • A provider-neutral request format — the gateway translates.
  • Visibility — log served_by_tier on every trace; alert when fallback traffic rises.

A real-life example

At 8:05 pm on the first night of a Diwali sale, a fintech support bot's primary provider starts returning 429s and slow 5xx errors for 40% of calls.

  • Retries with jitter absorb the first minute, but p95 latency rises to 9 s.
  • The circuit breaker opens for the primary region after 20 failures in 30 s; traffic moves to the same model family on a second cloud provider, where the team had reserved quota and validated the prompt a month earlier.
  • When that also hits its token limit at 8:40 pm, "where is my order" questions — 45% of traffic — go to the deterministic order-status API with a template answer, and other questions go to a smaller self-hosted model with a banner saying answers may be brief.

The dashboard shows 99.7% success, but also 31% of traffic served by tiers 3–5. The post-incident review uses that number to justify larger reserved quota for the next sale.

Follow-up questions to expect

  • "Which errors should never be retried?" — 400-level request errors (bad input, context too long, content policy); fix or route them instead.
  • "How do you avoid retries making an outage worse?" — Jitter, a small attempt cap, a retry budget, and a circuit breaker.
  • "How do you test fallbacks?" — Send real traffic (even 1%) to the fallback all the time, and run game days that disable the primary.