Live Coding Interview Prep

Course Content

Live Coding Interview Prep

7 sections · 50 lessons

Implement retry logic for failed tool calls.


What you need to know

Idempotent means doing the same call twice has the same effect as once. Reading an order is idempotent; charging a card is not. Retrying a non-idempotent call after a timeout can charge the customer twice, because the first call may have succeeded and only the reply was lost. The safe fix is an idempotency key: a unique id sent with the request that the server uses to ignore duplicates.

Exponential backoff waits 0.5 s, 1 s, 2 s, 4 s… between attempts, so a struggling service gets room to recover. Jitter adds randomness so that a thousand clients that failed at the same moment do not all retry at the same moment too. "Full jitter" picks a random wait between 0 and the backoff value.

failureexamplewho handles it
transienttimeout, 429, 503your code, with backoff
bad argumentsunknown field, wrong typethe model, via an error observation
permanent403, account closednobody; report it and stop
Python
import random, timefrom collections.abc import Callableclass TransientError(Exception): ...class BadArgumentsError(Exception): ...class PermanentError(Exception): ...def call_tool(fn: Callable, args: dict, attempts: int = 3, base: float = 0.5,              cap: float = 8.0, idempotent: bool = True,              sleep: Callable[[float], None] = time.sleep) -> dict:    """Retry transient failures with full-jitter backoff; hand the rest to the model."""    last: Exception | None = None    for i in range(attempts):        try:            return {"ok": True, "result": fn(**args), "attempts": i + 1}        except (BadArgumentsError, TypeError) as exc:            return {"ok": False, "error": f"invalid arguments: {exc}", "model_can_fix": True}        except PermanentError as exc:            return {"ok": False, "error": str(exc), "model_can_fix": False}        except (TransientError, TimeoutError, ConnectionError) as exc:            last = exc            if not idempotent or i == attempts - 1:                break            sleep(random.uniform(0, min(cap, base * 2 ** i)))    # full jitter        except Exception as exc:                                  # unknown: do not retry            return {"ok": False, "error": f"{type(exc).__name__}: {exc}", "model_can_fix": False}    return {"ok": False, "error": f"failed after retries: {last}", "model_can_fix": False}

The tricky parts:

  • Order of except clauses. Bad arguments are caught first and returned immediately — retrying identical arguments can never succeed. TypeError is included because calling a function with a wrong keyword raises it.
  • The catch-all at the end does not retry. An exception you did not plan for is treated as permanent; retrying unknown errors is how infinite loops happen.
  • sleep is a parameter, so tests pass a fake that records delays instead of waiting.
  • No sleep after the last attempt (i == attempts - 1 breaks first). Sleeping and then giving up is wasted time.

Complexity: at most attempts calls. Worst-case waiting is the sum of the caps: with base=0.5 and 3 attempts, at most 0.5 + 1.0 = 1.5 seconds, and on average half that with full jitter. Space is O(1).

A real-life example

A flaky train-status API that times out twice, then answers:

Python
random.seed(7)delays: list[float] = []state = {"calls": 0}def pnr_status(pnr: str) -> str:    state["calls"] += 1    if state["calls"] < 3:        raise TimeoutError("upstream timeout")    return f"PNR {pnr}: confirmed, coach B2"print(call_tool(pnr_status, {"pnr": "4521337890"}, sleep=delays.append))# {'ok': True, 'result': 'PNR 4521337890: confirmed, coach B2', 'attempts': 3}print([round(d, 3) for d in delays])# [0.162, 0.151]print(call_tool(pnr_status, {"pnr_number": "4521337890"}, sleep=delays.append))# {'ok': False, 'error': "invalid arguments: pnr_status() got an unexpected keyword argument 'pnr_number'", 'model_can_fix': True}
attemptoutcomewait before next
1TimeoutErrorrandom in [0, 0.5] → 0.162 s
2TimeoutErrorrandom in [0, 1.0] → 0.151 s
3success–

The second call uses a wrong argument name. It fails on the first attempt with no waiting, and model_can_fix: True tells the agent loop to show the error to the model, which then retries with pnr.

Every railway, airline and payment integration behind an AI assistant needs this split, because those upstream APIs time out often and some of them move money.

Follow-up questions to expect

  • "The server sends Retry-After: 30 — what do you do?" — Honour it instead of your computed delay; it is the server telling you exactly when it will accept you. If it is longer than your request deadline, fail fast.
  • "How do you make a payment tool safe to retry?" — Generate an idempotency key once per logical action, send it on every attempt, and have the server return the first result for a repeated key.
  • "Can retries make an outage worse?" — Yes; that is a retry storm. Jitter, a low attempt count, and a circuit breaker that stops calling a failing service all limit it.