AutoGen Essentials

Course Content

AutoGen Essentials

7 sections · 28 lessons

How do you handle tool failures (timeouts, bad inputs, partial outputs) in an agent workflow?


What you need to know

What AutoGen does by default

In 0.4+, AutoGen validates the model's arguments against the tool's type hints. If validation fails or the function raises, it catches the error and returns a tool result marked is_error=True, with the exception text as content. The model then sees something like ValueError: invalid literal for int(). That is better than a crash, but it does not tell the model what to change.

Failure types and fixes

FailureWhere to fixHow
Bad arguments from the modelToolValidate with types or Pydantic; return what was wrong and a valid example
TimeoutToolasyncio.wait_for or client timeouts; return "timeout" plus a hint
Rate limit or 5xxToolRetry with exponential backoff and jitter (for example tenacity), 2 to 3 attempts
Partial dataToolReturn status: partial, counts and a cursor
Repeated failureAgent and teammax_tool_iterations, a failure counter, a "give up and report" instruction
Double side effects on retryToolIdempotency key on writes

A tool that fails well

Python
import asyncioasync def search_trains(src: str, dst: str, date: str) -> dict:    """Search trains between two station codes on a date (YYYY-MM-DD)."""    if len(src) != 3 or len(dst) != 3:        return {"ok": False, "error": "bad_station_code",                "hint": "Use 3-letter codes such as NDLS or SBC."}    try:        rows = await asyncio.wait_for(rail_api.search(src, dst, date), timeout=8)    except asyncio.TimeoutError:        return {"ok": False, "error": "timeout",                "hint": "Try once more later; do not change the dates."}    if len(rows) > 20:        return {"ok": True, "status": "partial", "returned": 20,                "total": len(rows), "trains": rows[:20]}    return {"ok": True, "status": "complete", "trains": rows}

(The station-code check is simplified; real codes can be 2 to 4 letters.) The model now gets an error name, a hint and a clear partial flag.

Retry in code, not in conversation

A retry inside the tool costs milliseconds. A retry by the agent costs a full model call and can loop. Put backoff retries for transient errors in the tool, and let the agent see only the final outcome.

Define the give-up path

Put it in the system prompt: "If a tool fails twice, stop and tell the user what failed and what they can do." Pair it with a termination condition so the team ends. Log error rate per tool: a tool failing 30% of calls usually has a bad schema or description, not a bad model.

A real-life example

A travel-planning team booked trains through a partner API that timed out at peak hours around 10 am, when tatkal bookings open. The first version had no timeout: runs hung for 90 seconds, then the agent retried with slightly different dates, and three runs booked holds on the wrong day.

After the changes: an 8-second timeout, two backoff retries in the tool, an idempotency key on hold_seat, and the rule "on timeout, do not change dates". At peak, 7% of searches still time out, but the agent now replies "the rail system is busy, I'll retry in 2 minutes" and wrong-date holds dropped to zero.

Follow-up questions to expect

  • "Should the tool raise or return an error?" — Return a structured error for failures the model can react to. Raise only for bugs you want in your error tracker; AutoGen will still pass the text to the model.
  • "How do you stop the agent retrying forever?" — Cap max_tool_iterations, count failures in the tool and return "do not retry", and keep a token or message cap on the team.
  • "What about a tool that returns 50 of 4,000 rows?" — Say so in the result, with a cursor, so the agent does not summarise 50 rows as if they were all of them.