Course Content
AI Agent Fundamentals
5 sections · 13 lessons
Error Handling and Self-Correction Patterns
An analytics agent was asked for the Q3 churn rate. It ran this:
SELECT COUNT(*) FROM subscriptionsWHERE cancelled_at BETWEEN '2026-07-01' AND '2026-09-30';Result: 0. The agent retried. 0. It retried again with a two-second backoff. 0. Then it reported: "Q3 churn was 0%. No customers cancelled."
The real figure was 4.1%. The cancellations live in subscription_events with event_type = 'cancel'; the cancelled_at column on subscriptions was deprecated in March and is now always null.
Nothing failed. The database answered correctly every time — the query genuinely matches zero rows. The agent's retry machinery was working perfectly, and it was the wrong machinery entirely. Retrying assumes the action was right and the world was momentarily uncooperative. Here the action was wrong, and running it three more times could only produce the same wrong answer three more times.
Retry is the correct response to exactly one kind of error. Applied to the other three, it converts a recoverable mistake into a confident wrong answer.
Four kinds of error
| Kind | What broke | How you detect it | Signature |
|---|---|---|---|
| Execution | The tool could not run | Exception, non-2xx, timeout | Loud and immediate |
| Validation | The tool ran, the output is malformed | Schema check, parse failure, missing field | Loud, but only if you look |
| Logic | Everything ran; the plan was wrong | Only by checking the result against reality | Silent |
| State | The agent's beliefs no longer match the world | Later actions fail inexplicably, or corrupt data | Silent, then catastrophic |
Execution errors
Connection refused, HTTP 503, rate limit, timeout, permission denied. The system tells you plainly that the action did not happen. These are the easy ones — and precisely because they are easy, they are the only kind most error handlers are built for.
Validation errors
The call succeeded and returned rubbish: a JSON body that will not parse, a price field containing "N/A", an empty required array, a date of 0000-00-00. Detectable, but only if you validate. Code that does data["price"] * quantity without checking discovers the problem several steps later, in a stack trace that points nowhere near the cause.
Logic errors
Every call succeeded. Every response validated. The answer is wrong because the approach was wrong — the churn query above, or converting currency at last year's rate, or searching a knowledge base for a fact that only exists in a database. Nothing in the machinery can detect these, because from the machinery's point of view nothing went wrong.
State errors
The agent believes something that is no longer true. It cached the order status as "pending" at step 2, cancelled the order at step 5, and at step 9 tries to modify a shipping address on an order it cancelled itself. Or two agents both read a stock count of 3, both decrement, and the count ends at 2 instead of 1.
These are the worst kind because the failure surfaces far from its cause, and by then the agent has taken other actions premised on the false belief.
Matching the strategy to the error
| Error | Strategy | Because | Never do |
|---|---|---|---|
| Execution, transient | Retry with backoff | The action was right; the world was busy | Retry a write without an idempotency key |
| Execution, permanent | Fall back or escalate | Identical request, identical failure | Retry — it is pure waste |
| Validation | Reformulate the request | The tool worked; the input or expectation was off | Silently coerce and continue |
| Logic | Decompose and re-plan | The approach is wrong, not the execution | Retry — it cannot help |
| State | Refresh from source, then re-plan | Beliefs are stale; acting on them compounds the damage | Retry against the stale belief |
Strategy 1 — retry, for transient execution errors only
1TRANSIENT = {"connection_error", "connect_timeout", "rate_limited",2 "server_error_5xx", "timeout_unknown_outcome"}34def should_retry(result, tool_spec, attempt, max_attempts=4):5 if result["ok"] or attempt >= max_attempts:6 return False7 code = result["error"]["code"]8 if code not in TRANSIENT:9 return False10 if code == "timeout_unknown_outcome" and not tool_spec["idempotent"]:11 return False # read timeout on a write: outcome unknown12 return TrueStrategy 2 — fallback, when the primary route will not work
A fallback is a different way to get the same information, accepting a worse result rather than none.
1FALLBACKS = {2 "get_live_price": ["get_cached_price", "get_last_close"],3 "search_semantic": ["search_keyword"],4 "geocode_precise": ["geocode_city_centre"],5}67def with_fallback(tool, args, registry):8 result = execute(tool, args)9 if result["ok"]:10 return result1112 for alt in FALLBACKS.get(tool, []):13 alt_result = execute(alt, adapt_args(tool, alt, args))14 if alt_result["ok"]:15 alt_result["meta"]["degraded"] = True16 alt_result["meta"]["note"] = (17 f"{tool} unavailable ({result['error']['code']}); "18 f"used {alt} instead. This result may be less "19 f"accurate or out of date.")20 return alt_result2122 return resultThe degraded flag is the part that matters. An agent that quietly substitutes yesterday's closing price for a live quote and reports it as live is worse than one that fails, because the failure is now invisible. Every fallback must announce itself in the observation so the model can qualify its answer.
Strategy 3 — decompose, for logic errors
When the plan is wrong, no amount of re-execution helps. The agent must break the goal into smaller pieces and verify each one. Applied to the churn case:
FAILED PLAN count rows in subscriptions where cancelled_at is in Q3 -> 0, which is implausible for 12,400 subscribersDECOMPOSED 1. Does the table have any non-null cancelled_at at all? SELECT COUNT(*) FROM subscriptions WHERE cancelled_at IS NOT NULL; -> 0 ==> the column is unused. The plan's premise is false. 2. Where are cancellations actually recorded? describe_schema("subscription%") ==> subscription_events(subscription_id, event_type, occurred_at) 3. Count cancels in Q3 from the real source. SELECT COUNT(DISTINCT subscription_id) FROM subscription_events WHERE event_type = 'cancel' AND occurred_at >= '2026-07-01' AND occurred_at < '2026-10-01'; -> 508 4. Denominator: active at the start of Q3. -> 12400 5. 508 / 12400 = 0.04097 -> 4.10%Step 1 is the move worth learning. Before assuming a zero is real, test whether the mechanism that would produce a non-zero even exists. COUNT(*) WHERE cancelled_at IS NOT NULL returning 0 across the whole table is decisive: it is not that Q3 was quiet, it is that the column is dead.
Check the arithmetic: 508 ÷ 12,400 = 0.040967…, which is 4.10% to two decimal places. Against a reported 0%, the error was not small — it was the entire quantity.
The general trigger: a plausible-looking result that is implausible in context. Zero cancellations from 12,400 subscribers. A 30-second flight. A refund larger than the order. Build these sanity bounds into your tools, because they are the only automatic detector logic errors have:
1def sanity_check(metric, value, context):2 if metric == "churn_rate":3 if value == 0 and context["subscriber_count"] > 1000:4 return ("Churn of exactly 0 across "5 f"{context['subscriber_count']} subscribers is "6 "implausible. Verify the data source before "7 "reporting - the column may be unpopulated.")8 if value > 0.5:9 return (f"Churn of {value:.0%} is extreme. Check whether "10 "the denominator is correct.")11 return NoneState errors — restore, do not retry
1def refresh_and_replan(state, entity_type, entity_id, reason):2 truth = load_from_source(entity_type, entity_id) # authoritative3 stale = {k: (state.get(k), truth.get(k))4 for k in truth if state.get(k) != truth.get(k)}56 new_state = {**state, **truth}7 new_state["invalidated"] = list(stale)8 observation = (9 f"State was stale ({reason}). Refreshed {entity_type} "10 f"{entity_id} from source. Changed: "11 + "; ".join(f"{k}: {old!r} -> {new!r}"12 for k, (old, new) in stale.items())13 + ". Any plan step that assumed the old values must be redone."14 )15 return new_state, observationTwo things happen there. The state is replaced from the authoritative source rather than patched. And the agent is told which fields changed, so it can work out which of its earlier conclusions are now void. An observation reading only "state refreshed" leaves the model with no way to know what to reconsider.
Putting it together
1def handle(tool, args, state, registry, attempt=1):2 result = execute(tool, args)34 # 1. Execution errors5 if not result["ok"]:6 code = result["error"]["code"]7 if should_retry(result, registry[tool], attempt):8 time.sleep(backoff_delay(attempt))9 return handle(tool, args, state, registry, attempt + 1)10 if FALLBACKS.get(tool):11 alt = with_fallback(tool, args, registry)12 if alt["ok"]:13 return alt, state14 if code in ("stale_state", "conflict_409"):15 state, obs = refresh_and_replan(16 state, registry[tool]["entity"], args.get("id"), code)17 return err(code, obs), state18 return result, state1920 # 2. Validation errors21 problem = validate_output(result["data"], registry[tool]["returns"])22 if problem:23 return err("invalid_output",24 f"{tool} returned data failing its contract: "25 f"{problem}. Do not use this value. Try different "26 f"arguments or a different tool."), state2728 # 3. Logic errors - heuristic only, but better than nothing29 warning = sanity_check(registry[tool].get("metric"),30 result["data"], state)31 if warning:32 result["meta"]["sanity_warning"] = warning3334 return result, apply_result(state, tool, args, result)The ordering is the design. Execution errors first because nothing else is meaningful if the call did not happen. Validation next because a malformed result must never reach reasoning. Sanity checks last, attached as a warning rather than a failure — logic errors cannot be proven automatically, only flagged for the model's attention.
Self-correction patterns
Pattern 1 — reasoning-trace review
Before finalising, the agent re-reads its own trace looking for specific defects. The prompt must name what to look for; a generic "check your work" produces a generic "looks correct".
1REVIEW = """Review the trace below for these specific defects.2For each, answer YES or NO and cite the step number.341. Does any claim in the final answer rest on data that was never5 actually observed in a tool result?62. Was any tool result with 0 rows, an empty list, or a null7 treated as a meaningful finding rather than a possible error?83. Was any number computed mentally rather than by a tool?94. Was any tool called with arguments contradicted by an earlier10 observation?115. Does the answer address the question that was asked, or a12 related one?1314Trace:15{trace}1617Final answer:18{answer}"""Question 2 is the one that catches the churn failure. It converts an unarticulated assumption — "0 is the answer" — into an explicit yes/no the model must commit to.
Pattern 2 — output validation
Mechanical checks against the goal's requirements. This is the highest-value pattern, because it does not rely on the model's judgement at all.
1def validate_answer(answer, requirements, trace):2 problems = []34 for number in extract_numbers(answer):5 if not appears_in_observations(number, trace):6 problems.append(7 f"The figure {number} does not appear in any tool "8 f"result. Either cite the observation it came from "9 f"or remove it.")1011 if requirements.get("min_sources"):12 n = len(extract_urls(answer))13 if n < requirements["min_sources"]:14 problems.append(15 f"{n} sources cited; {requirements['min_sources']} "16 f"required.")1718 if requirements.get("must_include"):19 for field in requirements["must_include"]:20 if field not in answer:21 problems.append(f"Answer omits required field "22 f"'{field}'.")23 return problemsThe first check is the strongest anti-hallucination device available to an agent, and it is pure string matching. Every number in the answer must trace to an observation. A model cannot argue with a grep.
Pattern 3 — independent verification
A second model call, given the answer but not the reasoning, tries to reach the same conclusion or to falsify it. Withholding the trace is essential — a verifier shown the original reasoning tends to agree with it.
The arithmetic on this is instructive. Suppose a single checker independently catches 70% of errors. Two independent checkers miss an error only if both miss it: 0.30 × 0.30 = 0.09, so they catch 91%. Three catch 97.3%.
That is the ideal case, and it does not survive contact with reality. If both checkers are the same model with the same training and the same blind spots, their misses are strongly correlated. When correlation is high, the second checker's catch rate on the errors the first missed might be 20% rather than 70% — giving 0.30 × 0.80 = 0.24 missed, so 76% caught rather than 91%. You paid double for a six-point gain.
| Verifier | Cost | Catches | Independent of the agent? |
|---|---|---|---|
| Unit tests / schema check | Near zero | Validation errors, contract breaks | Completely |
| Sanity bounds on values | Near zero | Implausible logic errors | Completely |
| Number-provenance grep | Near zero | Fabricated figures | Completely |
| Re-query via a different tool | One tool call | Wrong data source, stale state | Largely |
| Same model, sees the trace | One model call | Little — it agrees with itself | No |
| Same model, answer only | One model call | Some reasoning errors | Partly |
| Different model, answer only | One model call | More, blind spots differ | Mostly |
| Human review | Minutes | Nearly everything | Yes |
The limits
Self-correction works when there is an external signal — a test that fails, a schema that rejects, a second data source that disagrees, a bound that is violated. It works poorly when the only signal is the model's own judgement, because the reasoning being evaluated and the reasoning doing the evaluating are the same reasoning. A model confident in a wrong answer will be confident in its review of that answer.
Self-correction is not the model checking itself. It is the model checking itself against something. No external referent, no correction.
Two further failure modes worth naming. Over-correction: an agent told to review its work will find something to change even when the answer was right, because "no problems found" reads as a failure to do the task. Give the review a defect list to test against and permission to return an empty list. Correction loops: the agent revises, reviews the revision, revises again, indefinitely. Cap it at one or two rounds; the third round of self-review almost never improves anything.
Where people get this wrong
One handler for all four error types. A single except Exception: retry(3) is right for one class and actively harmful for three. It is the entire churn incident in one line.
Treating empty results as answers. Zero rows, an empty list, a null. These are as likely to mean "your query was wrong" as "there is nothing there", and an agent cannot tell the difference without checking. Make your tools distinguish: {"rows": 0, "note": "the filtered column contains no non-null values in this table at all"} is a different observation from {"rows": 0}.
Fallbacks that hide themselves. Degrading silently converts a visible failure into an invisible wrong answer.
Asking the model to "double-check". Without a defect list it produces reassurance, not review.
Showing the verifier the reasoning. Contaminates the check. Give it the answer and the raw observations; withhold the chain that produced it.
Unlimited correction rounds. Each round costs a model call and, past the second, changes correct things into incorrect ones roughly as often as the reverse.
What this means when you build one
Classify before you react. Every error path in your agent should first answer "which of the four is this?" and only then choose a response. That single branch is worth more than any amount of retry tuning, because the expensive failures are the ones where retry was never the right tool.
Then invest disproportionately in mechanical verification. Sanity bounds on every numeric output, a provenance check that every figure in the answer appears in some observation, schema validation on every tool return. These cost microseconds, need no model call, and have no blind spots that correlate with the agent's. One grep that catches fabricated numbers is worth more than three rounds of the model reviewing itself.
Make degradation loud. Every fallback, every cached value, every truncated result must say so in the observation, so an answer built on weak foundations is qualified rather than confident.
And build the habit that would have caught the churn failure: when a result is surprising, verify the mechanism before believing the number. Zero cancellations from 12,400 subscribers is not a finding — it is a question. The agent that asks "is this column ever populated?" gets 4.1%. The agent that retries three times gets 0%, and gets it faster, and gets it wrong.