Course Content
Live Coding Interview Prep
7 sections · 50 lessons
Implement exponential backoff retry for API calls.
What you need to know
Transient errors go away if you wait: rate limits (429), overloaded servers (503, 529), timeouts. Permanent errors do not: a bad request (400), a wrong API key (401), a missing model (404). Retrying permanent errors multiplies latency by the attempt count and changes nothing.
Exponential backoff doubles the wait each time: 0.5 s, 1 s, 2 s, 4 s. The service gets progressively more room to recover.
Jitter randomises the wait. If 5,000 clients hit a 429 at 12:00:00 and all wait exactly 1 second, they all retry at 12:00:01 and cause the same spike. With full jitter each waits a random time between 0 and the backoff value, spreading the retries out.
Retry-After is a response header where the server says how many seconds to wait. It beats your guess.
SDKs already retry. The Anthropic Python SDK retries 408, 409, 429, 5xx and connection errors twice by default (max_retries). If you add your own retry layer, set max_retries=0 on the client, or the two loops multiply: 3 SDK attempts × 5 of yours = 15 calls.
1import functools, random, time2from collections.abc import Callable34RETRYABLE_STATUS = {408, 409, 429, 500, 502, 503, 504, 529}56def _status(exc: Exception) -> int | None:7 return getattr(exc, "status_code", None)89def _retry_after(exc: Exception) -> float | None:10 headers = getattr(getattr(exc, "response", None), "headers", None) or {}11 value = headers.get("retry-after")12 try:13 return float(value) if value is not None else None14 except ValueError:15 return None # an HTTP-date; fall back to backoff1617def retry_with_backoff(attempts: int = 5, base: float = 0.5, cap: float = 30.0,18 max_wait: float = 60.0, sleep: Callable[[float], None] = time.sleep):19 """Retry transient failures with full-jitter exponential backoff."""20 def decorator(fn):21 @functools.wraps(fn)22 def wrapper(*args, **kwargs):23 for i in range(attempts):24 try:25 return fn(*args, **kwargs)26 except Exception as exc:27 retryable = _status(exc) in RETRYABLE_STATUS or isinstance(28 exc, (TimeoutError, ConnectionError))29 if not retryable or i == attempts - 1:30 raise31 delay = _retry_after(exc)32 if delay is None:33 delay = random.uniform(0, min(cap, base * 2 ** i))34 if delay > max_wait:35 raise # waiting that long breaks our own deadline36 sleep(delay)37 return wrapper38 return decoratorUsing it with the SDK's own retries switched off:
1from anthropic import Anthropic23client = Anthropic(max_retries=0)45@retry_with_backoff(attempts=4)6def call_model(prompt: str):7 return client.messages.create(model="claude-opus-5", max_tokens=16000,8 messages=[{"role": "user", "content": prompt}])The tricky parts:
raisewith no argument re-raises the original exception with its traceback — the caller sees the real 400, not a wrapper.- The last attempt raises instead of sleeping. Sleeping and then giving up is pure waste.
Retry-Aftercan be a date, not a number. TheValueErrorbranch falls back to computed backoff rather than crashing inside the error handler.sleepis injectable, so a test can record the delays instead of waiting for them.
Complexity: at most attempts calls and attempts − 1 sleeps. With base=0.5, cap=30 and 5 attempts, the caps are 0.5 + 1 + 2 + 4 = 7.5 seconds of waiting at most, 3.75 seconds on average with full jitter. Space O(1).
A real-life example
A fake API that returns 429 twice (the second time with Retry-After: 2), then succeeds; then one that returns 400:
1class APIError(Exception):2 def __init__(self, status_code, headers=None):3 super().__init__(f"HTTP {status_code}")4 self.status_code = status_code5 self.response = type("R", (), {"headers": headers or {}})()67random.seed(1)8waits: list[float] = []9replies = iter([APIError(429), APIError(429, {"retry-after": "2"}), "ok"])1011@retry_with_backoff(attempts=5, sleep=waits.append)12def flaky():13 r = next(replies)14 if isinstance(r, Exception):15 raise r16 return r1718print(flaky(), [round(w, 3) for w in waits]) # ok [0.067, 2.0]1920@retry_with_backoff(sleep=waits.append)21def bad_request():22 raise APIError(400)23try:24 bad_request()25except APIError as e:26 print(e, len(waits)) # HTTP 400 2 (no new waits)| attempt | result | wait chosen | why |
|---|---|---|---|
| 1 | 429 | 0.067 s | random in [0, 0.5] |
| 2 | 429 + Retry-After: 2 | 2.0 s | server's value wins |
| 3 | ok | – | returned |
The 400 is raised on the first attempt: no sleep, no retry.
During a big sale, an e-commerce catalogue job that enriches 200,000 product listings with an LLM hits rate limits constantly; jittered backoff is what lets it finish without hammering the API.
Follow-up questions to expect
- "What is a retry budget?" — A cap on retries as a share of total traffic (say 10%). When the service is really down, you stop retrying altogether instead of tripling the load on it.
- "Is retrying a POST safe?" — Only if it is idempotent or carries an idempotency key. An LLM completion is safe to repeat (it costs money but changes nothing); creating an order is not.
- "Would you use a library?" — Yes:
tenacityin Python, or the SDK's built-in retries. Writing it by hand in an interview shows you know what those libraries decide for you.