Course Content
Live Coding Interview Prep
7 sections · 50 lessons
Implement retry logic for failed tool calls.
What you need to know
Idempotent means doing the same call twice has the same effect as once. Reading an order is idempotent; charging a card is not. Retrying a non-idempotent call after a timeout can charge the customer twice, because the first call may have succeeded and only the reply was lost. The safe fix is an idempotency key: a unique id sent with the request that the server uses to ignore duplicates.
Exponential backoff waits 0.5 s, 1 s, 2 s, 4 s… between attempts, so a struggling service gets room to recover. Jitter adds randomness so that a thousand clients that failed at the same moment do not all retry at the same moment too. "Full jitter" picks a random wait between 0 and the backoff value.
| failure | example | who handles it |
|---|---|---|
| transient | timeout, 429, 503 | your code, with backoff |
| bad arguments | unknown field, wrong type | the model, via an error observation |
| permanent | 403, account closed | nobody; report it and stop |
1import random, time2from collections.abc import Callable34class TransientError(Exception): ...5class BadArgumentsError(Exception): ...6class PermanentError(Exception): ...78def call_tool(fn: Callable, args: dict, attempts: int = 3, base: float = 0.5,9 cap: float = 8.0, idempotent: bool = True,10 sleep: Callable[[float], None] = time.sleep) -> dict:11 """Retry transient failures with full-jitter backoff; hand the rest to the model."""12 last: Exception | None = None13 for i in range(attempts):14 try:15 return {"ok": True, "result": fn(**args), "attempts": i + 1}16 except (BadArgumentsError, TypeError) as exc:17 return {"ok": False, "error": f"invalid arguments: {exc}", "model_can_fix": True}18 except PermanentError as exc:19 return {"ok": False, "error": str(exc), "model_can_fix": False}20 except (TransientError, TimeoutError, ConnectionError) as exc:21 last = exc22 if not idempotent or i == attempts - 1:23 break24 sleep(random.uniform(0, min(cap, base * 2 ** i))) # full jitter25 except Exception as exc: # unknown: do not retry26 return {"ok": False, "error": f"{type(exc).__name__}: {exc}", "model_can_fix": False}27 return {"ok": False, "error": f"failed after retries: {last}", "model_can_fix": False}The tricky parts:
- Order of
exceptclauses. Bad arguments are caught first and returned immediately — retrying identical arguments can never succeed.TypeErroris included because calling a function with a wrong keyword raises it. - The catch-all at the end does not retry. An exception you did not plan for is treated as permanent; retrying unknown errors is how infinite loops happen.
sleepis a parameter, so tests pass a fake that records delays instead of waiting.- No sleep after the last attempt (
i == attempts - 1breaks first). Sleeping and then giving up is wasted time.
Complexity: at most attempts calls. Worst-case waiting is the sum of the caps: with base=0.5 and 3 attempts, at most 0.5 + 1.0 = 1.5 seconds, and on average half that with full jitter. Space is O(1).
A real-life example
A flaky train-status API that times out twice, then answers:
1random.seed(7)2delays: list[float] = []3state = {"calls": 0}4def pnr_status(pnr: str) -> str:5 state["calls"] += 16 if state["calls"] < 3:7 raise TimeoutError("upstream timeout")8 return f"PNR {pnr}: confirmed, coach B2"910print(call_tool(pnr_status, {"pnr": "4521337890"}, sleep=delays.append))11# {'ok': True, 'result': 'PNR 4521337890: confirmed, coach B2', 'attempts': 3}12print([round(d, 3) for d in delays])13# [0.162, 0.151]14print(call_tool(pnr_status, {"pnr_number": "4521337890"}, sleep=delays.append))15# {'ok': False, 'error': "invalid arguments: pnr_status() got an unexpected keyword argument 'pnr_number'", 'model_can_fix': True}| attempt | outcome | wait before next |
|---|---|---|
| 1 | TimeoutError | random in [0, 0.5] → 0.162 s |
| 2 | TimeoutError | random in [0, 1.0] → 0.151 s |
| 3 | success | – |
The second call uses a wrong argument name. It fails on the first attempt with no waiting, and model_can_fix: True tells the agent loop to show the error to the model, which then retries with pnr.
Every railway, airline and payment integration behind an AI assistant needs this split, because those upstream APIs time out often and some of them move money.
Follow-up questions to expect
- "The server sends
Retry-After: 30— what do you do?" — Honour it instead of your computed delay; it is the server telling you exactly when it will accept you. If it is longer than your request deadline, fail fast. - "How do you make a payment tool safe to retry?" — Generate an idempotency key once per logical action, send it on every attempt, and have the server return the first result for a repeated key.
- "Can retries make an outage worse?" — Yes; that is a retry storm. Jitter, a low attempt count, and a circuit breaker that stops calling a failing service all limit it.