AI Agent Fundamentals

Error Handling and Self-Correction Patterns


An analytics agent was asked for the Q3 churn rate. It ran this:

SQL
SELECT COUNT(*) FROM subscriptionsWHERE cancelled_at BETWEEN '2026-07-01' AND '2026-09-30';

Result: 0. The agent retried. 0. It retried again with a two-second backoff. 0. Then it reported: "Q3 churn was 0%. No customers cancelled."

The real figure was 4.1%. The cancellations live in subscription_events with event_type = 'cancel'; the cancelled_at column on subscriptions was deprecated in March and is now always null.

Nothing failed. The database answered correctly every time — the query genuinely matches zero rows. The agent's retry machinery was working perfectly, and it was the wrong machinery entirely. Retrying assumes the action was right and the world was momentarily uncooperative. Here the action was wrong, and running it three more times could only produce the same wrong answer three more times.

Retry is the correct response to exactly one kind of error. Applied to the other three, it converts a recoverable mistake into a confident wrong answer.

The error decides the strategyerrortransient?wrong plan or state?retryfallbackdecomposerestore state
Retrying a logic error just runs the same wrong query again — only execution errors are worth a second attempt.

Four kinds of error

KindWhat brokeHow you detect itSignature
ExecutionThe tool could not runException, non-2xx, timeoutLoud and immediate
ValidationThe tool ran, the output is malformedSchema check, parse failure, missing fieldLoud, but only if you look
LogicEverything ran; the plan was wrongOnly by checking the result against realitySilent
StateThe agent's beliefs no longer match the worldLater actions fail inexplicably, or corrupt dataSilent, then catastrophic

Execution errors

Connection refused, HTTP 503, rate limit, timeout, permission denied. The system tells you plainly that the action did not happen. These are the easy ones — and precisely because they are easy, they are the only kind most error handlers are built for.

Validation errors

The call succeeded and returned rubbish: a JSON body that will not parse, a price field containing "N/A", an empty required array, a date of 0000-00-00. Detectable, but only if you validate. Code that does data["price"] * quantity without checking discovers the problem several steps later, in a stack trace that points nowhere near the cause.

Logic errors

Every call succeeded. Every response validated. The answer is wrong because the approach was wrong — the churn query above, or converting currency at last year's rate, or searching a knowledge base for a fact that only exists in a database. Nothing in the machinery can detect these, because from the machinery's point of view nothing went wrong.

State errors

The agent believes something that is no longer true. It cached the order status as "pending" at step 2, cancelled the order at step 5, and at step 9 tries to modify a shipping address on an order it cancelled itself. Or two agents both read a stock count of 3, both decrement, and the count ends at 2 instead of 1.

These are the worst kind because the failure surfaces far from its cause, and by then the agent has taken other actions premised on the false belief.

Matching the strategy to the error

ErrorStrategyBecauseNever do
Execution, transientRetry with backoffThe action was right; the world was busyRetry a write without an idempotency key
Execution, permanentFall back or escalateIdentical request, identical failureRetry — it is pure waste
ValidationReformulate the requestThe tool worked; the input or expectation was offSilently coerce and continue
LogicDecompose and re-planThe approach is wrong, not the executionRetry — it cannot help
StateRefresh from source, then re-planBeliefs are stale; acting on them compounds the damageRetry against the stale belief

Strategy 1 — retry, for transient execution errors only

Python
TRANSIENT = {"connection_error", "connect_timeout", "rate_limited",             "server_error_5xx", "timeout_unknown_outcome"}def should_retry(result, tool_spec, attempt, max_attempts=4):    if result["ok"] or attempt >= max_attempts:        return False    code = result["error"]["code"]    if code not in TRANSIENT:        return False    if code == "timeout_unknown_outcome" and not tool_spec["idempotent"]:        return False              # read timeout on a write: outcome unknown    return True

Strategy 2 — fallback, when the primary route will not work

A fallback is a different way to get the same information, accepting a worse result rather than none.

Python
FALLBACKS = {    "get_live_price":   ["get_cached_price", "get_last_close"],    "search_semantic":  ["search_keyword"],    "geocode_precise":  ["geocode_city_centre"],}def with_fallback(tool, args, registry):    result = execute(tool, args)    if result["ok"]:        return result    for alt in FALLBACKS.get(tool, []):        alt_result = execute(alt, adapt_args(tool, alt, args))        if alt_result["ok"]:            alt_result["meta"]["degraded"] = True            alt_result["meta"]["note"] = (                f"{tool} unavailable ({result['error']['code']}); "                f"used {alt} instead. This result may be less "                f"accurate or out of date.")            return alt_result    return result

The degraded flag is the part that matters. An agent that quietly substitutes yesterday's closing price for a live quote and reports it as live is worse than one that fails, because the failure is now invisible. Every fallback must announce itself in the observation so the model can qualify its answer.

Strategy 3 — decompose, for logic errors

When the plan is wrong, no amount of re-execution helps. The agent must break the goal into smaller pieces and verify each one. Applied to the churn case:

Text
FAILED PLAN  count rows in subscriptions where cancelled_at is in Q3  -> 0, which is implausible for 12,400 subscribersDECOMPOSED  1. Does the table have any non-null cancelled_at at all?     SELECT COUNT(*) FROM subscriptions     WHERE cancelled_at IS NOT NULL;              -> 0     ==> the column is unused. The plan's premise is false.  2. Where are cancellations actually recorded?     describe_schema("subscription%")     ==> subscription_events(subscription_id, event_type, occurred_at)  3. Count cancels in Q3 from the real source.     SELECT COUNT(DISTINCT subscription_id)     FROM subscription_events     WHERE event_type = 'cancel'       AND occurred_at >= '2026-07-01'       AND occurred_at <  '2026-10-01';           -> 508  4. Denominator: active at the start of Q3.               -> 12400  5. 508 / 12400 = 0.04097 -> 4.10%

Step 1 is the move worth learning. Before assuming a zero is real, test whether the mechanism that would produce a non-zero even exists. COUNT(*) WHERE cancelled_at IS NOT NULL returning 0 across the whole table is decisive: it is not that Q3 was quiet, it is that the column is dead.

Check the arithmetic: 508 ÷ 12,400 = 0.040967…, which is 4.10% to two decimal places. Against a reported 0%, the error was not small — it was the entire quantity.

The general trigger: a plausible-looking result that is implausible in context. Zero cancellations from 12,400 subscribers. A 30-second flight. A refund larger than the order. Build these sanity bounds into your tools, because they are the only automatic detector logic errors have:

Python
def sanity_check(metric, value, context):    if metric == "churn_rate":        if value == 0 and context["subscriber_count"] > 1000:            return ("Churn of exactly 0 across "                    f"{context['subscriber_count']} subscribers is "                    "implausible. Verify the data source before "                    "reporting - the column may be unpopulated.")        if value > 0.5:            return (f"Churn of {value:.0%} is extreme. Check whether "                    "the denominator is correct.")    return None

State errors — restore, do not retry

Python
def refresh_and_replan(state, entity_type, entity_id, reason):    truth = load_from_source(entity_type, entity_id)   # authoritative    stale = {k: (state.get(k), truth.get(k))             for k in truth if state.get(k) != truth.get(k)}    new_state = {**state, **truth}    new_state["invalidated"] = list(stale)    observation = (        f"State was stale ({reason}). Refreshed {entity_type} "        f"{entity_id} from source. Changed: "        + "; ".join(f"{k}: {old!r} -> {new!r}"                    for k, (old, new) in stale.items())        + ". Any plan step that assumed the old values must be redone."    )    return new_state, observation

Two things happen there. The state is replaced from the authoritative source rather than patched. And the agent is told which fields changed, so it can work out which of its earlier conclusions are now void. An observation reading only "state refreshed" leaves the model with no way to know what to reconsider.

Putting it together

Python
def handle(tool, args, state, registry, attempt=1):    result = execute(tool, args)    # 1. Execution errors    if not result["ok"]:        code = result["error"]["code"]        if should_retry(result, registry[tool], attempt):            time.sleep(backoff_delay(attempt))            return handle(tool, args, state, registry, attempt + 1)        if FALLBACKS.get(tool):            alt = with_fallback(tool, args, registry)            if alt["ok"]:                return alt, state        if code in ("stale_state", "conflict_409"):            state, obs = refresh_and_replan(                state, registry[tool]["entity"], args.get("id"), code)            return err(code, obs), state        return result, state    # 2. Validation errors    problem = validate_output(result["data"], registry[tool]["returns"])    if problem:        return err("invalid_output",                   f"{tool} returned data failing its contract: "                   f"{problem}. Do not use this value. Try different "                   f"arguments or a different tool."), state    # 3. Logic errors - heuristic only, but better than nothing    warning = sanity_check(registry[tool].get("metric"),                           result["data"], state)    if warning:        result["meta"]["sanity_warning"] = warning    return result, apply_result(state, tool, args, result)

The ordering is the design. Execution errors first because nothing else is meaningful if the call did not happen. Validation next because a malformed result must never reach reasoning. Sanity checks last, attached as a warning rather than a failure — logic errors cannot be proven automatically, only flagged for the model's attention.

Self-correction patterns

Pattern 1 — reasoning-trace review

Before finalising, the agent re-reads its own trace looking for specific defects. The prompt must name what to look for; a generic "check your work" produces a generic "looks correct".

Python
REVIEW = """Review the trace below for these specific defects.For each, answer YES or NO and cite the step number.1. Does any claim in the final answer rest on data that was never   actually observed in a tool result?2. Was any tool result with 0 rows, an empty list, or a null   treated as a meaningful finding rather than a possible error?3. Was any number computed mentally rather than by a tool?4. Was any tool called with arguments contradicted by an earlier   observation?5. Does the answer address the question that was asked, or a   related one?Trace:{trace}Final answer:{answer}"""

Question 2 is the one that catches the churn failure. It converts an unarticulated assumption — "0 is the answer" — into an explicit yes/no the model must commit to.

Pattern 2 — output validation

Mechanical checks against the goal's requirements. This is the highest-value pattern, because it does not rely on the model's judgement at all.

Python
def validate_answer(answer, requirements, trace):    problems = []    for number in extract_numbers(answer):        if not appears_in_observations(number, trace):            problems.append(                f"The figure {number} does not appear in any tool "                f"result. Either cite the observation it came from "                f"or remove it.")    if requirements.get("min_sources"):        n = len(extract_urls(answer))        if n < requirements["min_sources"]:            problems.append(                f"{n} sources cited; {requirements['min_sources']} "                f"required.")    if requirements.get("must_include"):        for field in requirements["must_include"]:            if field not in answer:                problems.append(f"Answer omits required field "                                f"'{field}'.")    return problems

The first check is the strongest anti-hallucination device available to an agent, and it is pure string matching. Every number in the answer must trace to an observation. A model cannot argue with a grep.

Pattern 3 — independent verification

A second model call, given the answer but not the reasoning, tries to reach the same conclusion or to falsify it. Withholding the trace is essential — a verifier shown the original reasoning tends to agree with it.

The arithmetic on this is instructive. Suppose a single checker independently catches 70% of errors. Two independent checkers miss an error only if both miss it: 0.30 × 0.30 = 0.09, so they catch 91%. Three catch 97.3%.

That is the ideal case, and it does not survive contact with reality. If both checkers are the same model with the same training and the same blind spots, their misses are strongly correlated. When correlation is high, the second checker's catch rate on the errors the first missed might be 20% rather than 70% — giving 0.30 × 0.80 = 0.24 missed, so 76% caught rather than 91%. You paid double for a six-point gain.

VerifierCostCatchesIndependent of the agent?
Unit tests / schema checkNear zeroValidation errors, contract breaksCompletely
Sanity bounds on valuesNear zeroImplausible logic errorsCompletely
Number-provenance grepNear zeroFabricated figuresCompletely
Re-query via a different toolOne tool callWrong data source, stale stateLargely
Same model, sees the traceOne model callLittle — it agrees with itselfNo
Same model, answer onlyOne model callSome reasoning errorsPartly
Different model, answer onlyOne model callMore, blind spots differMostly
Human reviewMinutesNearly everythingYes

The limits

Self-correction works when there is an external signal — a test that fails, a schema that rejects, a second data source that disagrees, a bound that is violated. It works poorly when the only signal is the model's own judgement, because the reasoning being evaluated and the reasoning doing the evaluating are the same reasoning. A model confident in a wrong answer will be confident in its review of that answer.

Self-correction is not the model checking itself. It is the model checking itself against something. No external referent, no correction.

Two further failure modes worth naming. Over-correction: an agent told to review its work will find something to change even when the answer was right, because "no problems found" reads as a failure to do the task. Give the review a defect list to test against and permission to return an empty list. Correction loops: the agent revises, reviews the revision, revises again, indefinitely. Cap it at one or two rounds; the third round of self-review almost never improves anything.

Where people get this wrong

One handler for all four error types. A single except Exception: retry(3) is right for one class and actively harmful for three. It is the entire churn incident in one line.

Treating empty results as answers. Zero rows, an empty list, a null. These are as likely to mean "your query was wrong" as "there is nothing there", and an agent cannot tell the difference without checking. Make your tools distinguish: {"rows": 0, "note": "the filtered column contains no non-null values in this table at all"} is a different observation from {"rows": 0}.

Fallbacks that hide themselves. Degrading silently converts a visible failure into an invisible wrong answer.

Asking the model to "double-check". Without a defect list it produces reassurance, not review.

Showing the verifier the reasoning. Contaminates the check. Give it the answer and the raw observations; withhold the chain that produced it.

Unlimited correction rounds. Each round costs a model call and, past the second, changes correct things into incorrect ones roughly as often as the reverse.

What this means when you build one

Classify before you react. Every error path in your agent should first answer "which of the four is this?" and only then choose a response. That single branch is worth more than any amount of retry tuning, because the expensive failures are the ones where retry was never the right tool.

Then invest disproportionately in mechanical verification. Sanity bounds on every numeric output, a provenance check that every figure in the answer appears in some observation, schema validation on every tool return. These cost microseconds, need no model call, and have no blind spots that correlate with the agent's. One grep that catches fabricated numbers is worth more than three rounds of the model reviewing itself.

Make degradation loud. Every fallback, every cached value, every truncated result must say so in the observation, so an answer built on weak foundations is qualified rather than confident.

And build the habit that would have caught the churn failure: when a result is surprising, verify the mechanism before believing the number. Zero cancellations from 12,400 subscribers is not a finding — it is a question. The agent that asks "is this column ever populated?" gets 4.1%. The agent that retries three times gets 0%, and gets it faster, and gets it wrong.