Live Coding Interview Prep

Course Content

Live Coding Interview Prep

7 sections · 50 lessons

Implement exponential backoff retry for API calls.


What you need to know

Transient errors go away if you wait: rate limits (429), overloaded servers (503, 529), timeouts. Permanent errors do not: a bad request (400), a wrong API key (401), a missing model (404). Retrying permanent errors multiplies latency by the attempt count and changes nothing.

Exponential backoff doubles the wait each time: 0.5 s, 1 s, 2 s, 4 s. The service gets progressively more room to recover.

Jitter randomises the wait. If 5,000 clients hit a 429 at 12:00:00 and all wait exactly 1 second, they all retry at 12:00:01 and cause the same spike. With full jitter each waits a random time between 0 and the backoff value, spreading the retries out.

Retry-After is a response header where the server says how many seconds to wait. It beats your guess.

SDKs already retry. The Anthropic Python SDK retries 408, 409, 429, 5xx and connection errors twice by default (max_retries). If you add your own retry layer, set max_retries=0 on the client, or the two loops multiply: 3 SDK attempts × 5 of yours = 15 calls.

Python
import functools, random, timefrom collections.abc import CallableRETRYABLE_STATUS = {408, 409, 429, 500, 502, 503, 504, 529}def _status(exc: Exception) -> int | None:    return getattr(exc, "status_code", None)def _retry_after(exc: Exception) -> float | None:    headers = getattr(getattr(exc, "response", None), "headers", None) or {}    value = headers.get("retry-after")    try:        return float(value) if value is not None else None    except ValueError:        return None                                 # an HTTP-date; fall back to backoffdef retry_with_backoff(attempts: int = 5, base: float = 0.5, cap: float = 30.0,                       max_wait: float = 60.0, sleep: Callable[[float], None] = time.sleep):    """Retry transient failures with full-jitter exponential backoff."""    def decorator(fn):        @functools.wraps(fn)        def wrapper(*args, **kwargs):            for i in range(attempts):                try:                    return fn(*args, **kwargs)                except Exception as exc:                    retryable = _status(exc) in RETRYABLE_STATUS or isinstance(                        exc, (TimeoutError, ConnectionError))                    if not retryable or i == attempts - 1:                        raise                    delay = _retry_after(exc)                    if delay is None:                        delay = random.uniform(0, min(cap, base * 2 ** i))                    if delay > max_wait:                        raise                       # waiting that long breaks our own deadline                    sleep(delay)        return wrapper    return decorator

Using it with the SDK's own retries switched off:

Python
from anthropic import Anthropicclient = Anthropic(max_retries=0)@retry_with_backoff(attempts=4)def call_model(prompt: str):    return client.messages.create(model="claude-opus-5", max_tokens=16000,                                  messages=[{"role": "user", "content": prompt}])

The tricky parts:

  • raise with no argument re-raises the original exception with its traceback — the caller sees the real 400, not a wrapper.
  • The last attempt raises instead of sleeping. Sleeping and then giving up is pure waste.
  • Retry-After can be a date, not a number. The ValueError branch falls back to computed backoff rather than crashing inside the error handler.
  • sleep is injectable, so a test can record the delays instead of waiting for them.

Complexity: at most attempts calls and attempts − 1 sleeps. With base=0.5, cap=30 and 5 attempts, the caps are 0.5 + 1 + 2 + 4 = 7.5 seconds of waiting at most, 3.75 seconds on average with full jitter. Space O(1).

A real-life example

A fake API that returns 429 twice (the second time with Retry-After: 2), then succeeds; then one that returns 400:

Python
class APIError(Exception):    def __init__(self, status_code, headers=None):        super().__init__(f"HTTP {status_code}")        self.status_code = status_code        self.response = type("R", (), {"headers": headers or {}})()random.seed(1)waits: list[float] = []replies = iter([APIError(429), APIError(429, {"retry-after": "2"}), "ok"])@retry_with_backoff(attempts=5, sleep=waits.append)def flaky():    r = next(replies)    if isinstance(r, Exception):        raise r    return rprint(flaky(), [round(w, 3) for w in waits])      # ok [0.067, 2.0]@retry_with_backoff(sleep=waits.append)def bad_request():    raise APIError(400)try:    bad_request()except APIError as e:    print(e, len(waits))                            # HTTP 400 2  (no new waits)
attemptresultwait chosenwhy
14290.067 srandom in [0, 0.5]
2429 + Retry-After: 22.0 sserver's value wins
3ok–returned

The 400 is raised on the first attempt: no sleep, no retry.

During a big sale, an e-commerce catalogue job that enriches 200,000 product listings with an LLM hits rate limits constantly; jittered backoff is what lets it finish without hammering the API.

Follow-up questions to expect

  • "What is a retry budget?" — A cap on retries as a share of total traffic (say 10%). When the service is really down, you stop retrying altogether instead of tripling the load on it.
  • "Is retrying a POST safe?" — Only if it is idempotent or carries an idempotency key. An LLM completion is safe to repeat (it costs money but changes nothing); creating an order is not.
  • "Would you use a library?" — Yes: tenacity in Python, or the SDK's built-in retries. Writing it by hand in an interview shows you know what those libraries decide for you.