Course Content
AI Agent Fundamentals
5 sections · 13 lessons
Action Execution and Feedback Loops
A customer was refunded 4,250 rupees three times. The finance team found it four days later, 8,500 rupees down.
The log explains it in six lines:
11:04:02.118 POST /refunds {order: ORD-4471, paise: 425000}11:04:07.118 TIMEOUT after 5000ms - no response11:04:07.120 retry 1/311:04:12.120 TIMEOUT after 5000ms - no response11:04:12.121 retry 2/311:04:14.902 200 OK {refund_id: RF-99203, status: settled}The payment provider was slow that morning, not broken. All three requests arrived. All three were processed. The first two took 9 and 7 seconds to complete — longer than the 5-second client timeout, so the agent never saw the responses and assumed failure.
The retry logic was textbook. That is the problem: it was textbook retry logic applied to an action that must never be retried blindly. A timeout tells you that you stopped waiting. It tells you nothing about whether the other side did the work.
A timeout is not a failure. It is the absence of information. Treating it as a failure is how agents double-charge customers.
The execution pipeline
Between "the model asked for issue_refund" and "the observation goes back into the trace" sit seven stages. Skipping any of them produces a specific class of production incident.
model decision │ 1. VALIDATE shape, types, ranges, permission, ownership │ 2. EXECUTE with timeout, idempotency key, circuit breaker │ 3. RETRY? only if transient AND safe │ 4. NORMALISE one predictable result shape, always │ 5. FEEDBACK an observation the model can act on │ 6. MONITOR record latency, outcome, cost │ 7. UPDATE STATE atomically, or not at all │back into the loopStage 1 — validate before you execute
Model output is untrusted input. It is generated text, influenced by everything in the conversation including anything a user typed. Validate it as strictly as you would validate an HTTP request from a stranger.
1from decimal import Decimal23def validate(tool_name, args, context):4 spec = TOOLS.get(tool_name)5 if spec is None:6 return err("unknown_tool",7 f"No tool named {tool_name}. Available: "8 f"{', '.join(TOOLS)}")910 missing = [p for p in spec["required"] if p not in args]11 if missing:12 return err("missing_arguments",13 f"{tool_name} requires {missing}. "14 f"You supplied {sorted(args)}.")1516 extra = [k for k in args if k not in spec["properties"]]17 if extra:18 return err("unknown_arguments",19 f"{tool_name} has no parameters {extra}. "20 f"Valid: {sorted(spec['properties'])}.")2122 for key, value in args.items():23 want = spec["properties"][key]["type"]24 if not matches_type(value, want):25 return err("wrong_type",26 f"{key} must be {want}, got "27 f"{type(value).__name__} ({value!r}).")2829 if not context.user.can(spec["permission"]):30 return err("forbidden",31 f"{tool_name} requires {spec['permission']}, "32 f"which this session does not hold. "33 f"Escalate instead of retrying.")3435 if tool_name == "issue_refund":36 if Decimal(args["amount_paise"]) > context.max_refund_paise:37 return err("over_limit",38 f"Refund of {args['amount_paise']} paise exceeds "39 f"the {context.max_refund_paise} limit. Request "40 f"approval with request_approval().")41 if not owns_order(context.user, args["order_id"]):42 return err("not_owner",43 f"Order {args['order_id']} does not belong to "44 f"the customer in this conversation.")4546 return ok(args)Six checks, in a deliberate order: existence, completeness, no extras, types, permission, business rules. Cheapest and most fundamental first, so a typo never reaches an ownership query.
The extra check is the one people leave out and the one that catches hallucinated parameters. A model that emits {"order_id": "ORD-4471", "force": true} against a function with no force parameter would otherwise raise a TypeError deep in your call stack. Caught here, it becomes a readable observation.
The permission message is worth copying too: it names the missing permission and tells the agent not to retry. Without that instruction the agent will try the same call again, because from its side a failure looks like something worth another attempt.
Stage 2 — timeouts, done properly
Every external call needs a deadline. The naive implementation is worse than none:
1# WRONG - the thread keeps running, the work still happens2import threading34result = [None]5t = threading.Thread(target=lambda: result.__setitem__(0, call_api()))6t.start()7t.join(timeout=5)8if t.is_alive():9 raise TimeoutError() # thread continues; API still processingTwo things go wrong. The thread is not cancelled — it keeps running and keeps consuming a connection. And the remote side never learns you gave up, so it completes the work anyway. That is exactly the refund incident.
1import httpx, uuid23def execute(tool_name, args, *, timeout_s=5.0, idempotency_key=None):4 started = time.monotonic()5 try:6 resp = httpx.post(7 ENDPOINTS[tool_name],8 json=args,9 timeout=httpx.Timeout(connect=2.0, read=timeout_s,10 write=2.0, pool=1.0),11 headers={"Idempotency-Key": idempotency_key} if12 idempotency_key else {},13 )14 resp.raise_for_status()15 return ok(resp.json(), elapsed=time.monotonic() - started)1617 except httpx.ReadTimeout:18 return err("timeout_unknown_outcome",19 f"No response within {timeout_s}s. The operation MAY "20 f"have completed. Do not retry a write. Verify with "21 f"a read before taking further action.",22 elapsed=time.monotonic() - started)2324 except httpx.ConnectTimeout:25 return err("connect_timeout",26 "Could not open a connection. The request was never "27 "delivered, so it is safe to retry.",28 elapsed=time.monotonic() - started)Separating ConnectTimeout from ReadTimeout is the crux. A connect timeout means the request never arrived — retrying is provably safe. A read timeout means it arrived and you stopped listening — retrying is a coin flip. Most HTTP clients let you distinguish these, and most agent code throws the distinction away.
Idempotency keys
The proper fix for the whole class of problem: send a unique key with every write, and have the server return the original result for a repeat of the same key rather than doing the work twice.
1key = f"refund:{order_id}:{amount_paise}:{run_id}"23# attempt 1 -> server records key, processes, returns RF-992034# attempt 2 -> server sees key, returns RF-99203 without processing5# attempt 3 -> sameWith this in place the refund incident produces one refund and two identical responses. Build the key from the operation's identity, not from a random UUID per attempt — a fresh UUID on each retry defeats the entire mechanism, which is a surprisingly common bug.
Stage 3 — retry, only when it is safe
Retrying needs both conditions: the failure must be transient, and the action must be safe to repeat.
Transient or permanent
| Signal | Class | Retry? | Why |
|---|---|---|---|
| Connection refused / DNS failure | Transient | Yes | Never delivered; service may return |
| HTTP 429 rate limited | Transient | Yes — honour Retry-After | Explicitly "try later" |
| HTTP 500 / 502 / 503 | Transient | Yes | Server-side fault, often momentary |
| Read timeout | Unknown | Only if idempotent | Outcome genuinely undetermined |
| HTTP 400 bad request | Permanent | No | Identical request gives an identical error |
| HTTP 401 / 403 | Permanent | No | Credentials will not change on retry |
| HTTP 404 | Permanent | No | The resource does not exist |
| HTTP 409 conflict | Permanent | No — re-read first | State changed; the premise is stale |
| Validation error from your own code | Permanent | No — return to the model | Only different arguments can fix it |
Idempotent or not
| Action | Idempotent? | Retry policy |
|---|---|---|
get_order(id) | Yes | Retry freely |
search(query) | Yes | Retry freely |
set_status(id, "shipped") | Yes — same end state | Retry freely |
issue_refund(order, amount) | No | Only with an idempotency key |
send_email(to, body) | No | Only with an idempotency key |
append_row(table, data) | No | Only with an idempotency key |
increment_counter(id) | No | Never blind-retry |
Note set_status. It writes, and it is still idempotent, because running it twice leaves the world in the same state. Idempotence is about the end state, not about whether the operation mutates.
Exponential backoff with jitter
1import random, time23def with_backoff(fn, *, max_attempts=4, base_s=1.0, cap_s=30.0):4 for attempt in range(max_attempts):5 result = fn()6 if result["ok"] or not result.get("retryable"):7 return result8 if attempt == max_attempts - 1:9 return result1011 delay = min(base_s * (2 ** attempt), cap_s)12 delay = delay * (0.5 + random.random()) # jitter: 0.5x to 1.5x13 time.sleep(delay)14 return resultThe schedule with base_s = 1.0, before jitter:
| Attempt | Delay before it | Cumulative wait |
|---|---|---|
| 1 | 0s | 0s |
| 2 | 1s | 1s |
| 3 | 2s | 3s |
| 4 | 4s | 7s |
Four attempts cost about 7 seconds of waiting plus the request time (up to 10.5 seconds once jitter stretches the delays). Extend to six attempts and it is 1+2+4+8+16 = 31 seconds, which is usually longer than a user will tolerate. Four is a good default.
The jitter multiplier is not a nicety. Without it, 200 agents hitting the same rate limit all sleep exactly 1 second and then all retry in the same millisecond — a thundering herd that keeps the service down. Multiplying by a random factor between 0.5 and 1.5 spreads them out, and it is one line.
Circuit breakers
When a dependency is genuinely down, retrying every call makes things worse for everyone and burns your step budget. A breaker stops trying after a run of failures:
1class CircuitBreaker:2 def __init__(self, threshold=5, cooldown_s=60):3 self.threshold, self.cooldown = threshold, cooldown_s4 self.failures, self.opened_at = 0, None56 def allow(self):7 if self.opened_at is None:8 return True9 if time.monotonic() - self.opened_at > self.cooldown:10 self.opened_at, self.failures = None, 0 # half-open11 return True12 return False1314 def record(self, ok):15 if ok:16 self.failures, self.opened_at = 0, None17 else:18 self.failures += 119 if self.failures >= self.threshold:20 self.opened_at = time.monotonic()When the breaker is open, return an observation rather than an exception: "The payments service has failed 5 times in a row and is temporarily unavailable. Do not retry for 60 seconds. Use a different approach or escalate." That gives the agent a real decision to make.
Stage 4 — one predictable result shape
Every tool, on every path, returns the same envelope. No exceptions, no special cases.
1def ok(data, **meta):2 return {"ok": True, "data": data, "error": None, "meta": meta}34def err(code, message, **meta):5 return {"ok": False, "data": None,6 "error": {"code": code, "message": message}, "meta": meta}Why this rigidity pays for itself: your loop has exactly one place that branches on success, one place that formats the observation, and one place that records metrics. Without it, every tool needs bespoke handling, and the third-party tool someone adds next month will return a bare string and break the agent in a way nobody notices for a week.
Stage 5 — feedback the agent can use
An observation exists to help the model choose the next action. Judge every one by that standard.
| Useless | Useful | What the useful one enables |
|---|---|---|
Error | 404: no order 'ORD-447'. IDs are ORD- plus 5 digits; you supplied 4. | Fix the ID and retry |
[] | 0 results for "widget xl". The catalogue has 3 items matching "widget". Try a broader term. | Broaden instead of repeating |
Failed | Rate limited. Retry after 30s. 4 calls remain this minute. | Wait, or do something else meanwhile |
Success | Refund RF-99203 settled, 425000 paise. Funds visible in 3-5 working days. | Tell the customer something specific |
{...40KB of JSON...} | 12 orders (showing 5 newest). Total 51200 paise. Full list: use offset. | Answer without drowning in context |
Three rules generate all of these:
- Say what happened, specifically. Codes, counts, identifiers, amounts — not adjectives.
- Say what to do next. Retry, broaden, wait, escalate, stop. The model is choosing an action; give it a candidate.
- Say what you left out. Truncated, filtered, approximated, stale — anything the agent might otherwise mistake for the whole picture.
Write every observation as if it were the only sentence the model will read before its next decision. Because effectively, it is.
Stage 6 — monitoring
Record one row per action execution: run ID, step, tool, argument hash, outcome code, latency, retry count, tokens. Four aggregate signals then fall out of it.
| Signal | Healthy | Reading a bad value |
|---|---|---|
| Tool error rate | Under 5% | Above 10% means bad descriptions, not a bad model |
| p95 latency per tool | Well under the timeout | p95 near the timeout means timeouts are coming |
| Retry rate | Under 2% | High and rising means a dependency is degrading |
| Repeat-call rate | ≈ 0 | Anything above zero means uninformative observations |
Watch the relationship between p95 latency and your timeout in particular. If issue_refund normally completes in 800 ms with a p95 of 4.6 seconds and your timeout is 5 seconds, roughly 5% of calls are within 400 ms of timing out. That is not a hypothetical — that is the refund incident arriving next Tuesday.
Stage 7 — updating state atomically
The last trap. Consider:
1# WRONG - three chances to end up half-updated2state["refund_id"] = result["data"]["refund_id"]3state["balance_paise"] -= args["amount_paise"]4audit_log.write(result) # if this raises...5state["step"] += 1 # ...this never runsIf the audit write fails, state records a refund that no log knows about, and the step counter is wrong. Build the new state, then swap it in one assignment:
1def apply_result(state, tool, args, result):2 new = dict(state) # copy, do not mutate3 new["step"] = state["step"] + 14 new["history"] = state["history"] + [5 {"step": new["step"], "tool": tool, "args": args,6 "ok": result["ok"],7 "summary": summarise(result)}8 ]9 if result["ok"] and tool == "issue_refund":10 new["refund_id"] = result["data"]["refund_id"]11 new["balance_paise"] = (state["balance_paise"]12 - args["amount_paise"])13 return new # caller does state = newThree properties this buys. The old state is never damaged, so a failure mid-update leaves the agent in a known-good position. Every step lands in history whether it succeeded or not, so the trace is complete. And because the function is pure, you can replay a whole run from a list of results and get identical state — which is what makes debugging tractable.
Where people get this wrong
Retrying on timeout without an idempotency key. The incident at the top. If the action writes and you cannot prove the request never arrived, verify with a read before doing anything else.
One timeout for everything. A 30-second default because "some tools are slow" means a dead dependency wastes 30 seconds per attempt. Set it per tool from the observed p99, plus headroom.
Retrying permanent failures. Four attempts at an HTTP 400 cost four round trips and about 7 seconds to receive the same error. Consult the failure class before the retry count.
Fresh idempotency key per attempt. Generating a new UUID inside the retry loop makes every attempt look like a distinct operation. The key must be derived from the operation, not the attempt.
Logging only failures. You cannot compute a p95 latency or an error rate from failures alone. Record every execution.
Mutating state during execution. An exception halfway through leaves a half-applied update that no later code accounts for, and the resulting bugs are non-deterministic and miserable to chase.
What this means when you build one
Sort your tools into two lists on day one: those that are safe to repeat, and those that are not. Anything on the second list gets an idempotency key derived from the operation's identity before it ships. Not "later" — the incident that teaches you this lesson costs real money and real trust.
Set timeouts per tool from measured latency, and separate connect failures from read failures so your retry logic can tell "never arrived" from "outcome unknown". Those two situations demand opposite responses and most codebases collapse them into one.
Give every tool the same result envelope, and make error messages instructions rather than complaints. The single sentence "Do not retry a write; verify with a read first" in a timeout observation is worth more than a page of prompt engineering, because it reaches the model at the exact moment the decision is being made.
And keep state updates pure — build the new state, swap it once. When something eventually goes wrong at 2am, the difference between an agent whose entire history is replayable and one whose state was mutated in seven places is the difference between a ten-minute diagnosis and a lost afternoon.