Course Content
AutoGen Essentials
7 sections · 28 lessons
How do you handle tool failures (timeouts, bad inputs, partial outputs) in an agent workflow?
What you need to know
What AutoGen does by default
In 0.4+, AutoGen validates the model's arguments against the tool's type hints. If validation fails or the function raises, it catches the error and returns a tool result marked is_error=True, with the exception text as content. The model then sees something like ValueError: invalid literal for int(). That is better than a crash, but it does not tell the model what to change.
Failure types and fixes
| Failure | Where to fix | How |
|---|---|---|
| Bad arguments from the model | Tool | Validate with types or Pydantic; return what was wrong and a valid example |
| Timeout | Tool | asyncio.wait_for or client timeouts; return "timeout" plus a hint |
| Rate limit or 5xx | Tool | Retry with exponential backoff and jitter (for example tenacity), 2 to 3 attempts |
| Partial data | Tool | Return status: partial, counts and a cursor |
| Repeated failure | Agent and team | max_tool_iterations, a failure counter, a "give up and report" instruction |
| Double side effects on retry | Tool | Idempotency key on writes |
A tool that fails well
1import asyncio23async def search_trains(src: str, dst: str, date: str) -> dict:4 """Search trains between two station codes on a date (YYYY-MM-DD)."""5 if len(src) != 3 or len(dst) != 3:6 return {"ok": False, "error": "bad_station_code",7 "hint": "Use 3-letter codes such as NDLS or SBC."}8 try:9 rows = await asyncio.wait_for(rail_api.search(src, dst, date), timeout=8)10 except asyncio.TimeoutError:11 return {"ok": False, "error": "timeout",12 "hint": "Try once more later; do not change the dates."}13 if len(rows) > 20:14 return {"ok": True, "status": "partial", "returned": 20,15 "total": len(rows), "trains": rows[:20]}16 return {"ok": True, "status": "complete", "trains": rows}(The station-code check is simplified; real codes can be 2 to 4 letters.) The model now gets an error name, a hint and a clear partial flag.
Retry in code, not in conversation
A retry inside the tool costs milliseconds. A retry by the agent costs a full model call and can loop. Put backoff retries for transient errors in the tool, and let the agent see only the final outcome.
Define the give-up path
Put it in the system prompt: "If a tool fails twice, stop and tell the user what failed and what they can do." Pair it with a termination condition so the team ends. Log error rate per tool: a tool failing 30% of calls usually has a bad schema or description, not a bad model.
A real-life example
A travel-planning team booked trains through a partner API that timed out at peak hours around 10 am, when tatkal bookings open. The first version had no timeout: runs hung for 90 seconds, then the agent retried with slightly different dates, and three runs booked holds on the wrong day.
After the changes: an 8-second timeout, two backoff retries in the tool, an idempotency key on hold_seat, and the rule "on timeout, do not change dates". At peak, 7% of searches still time out, but the agent now replies "the rail system is busy, I'll retry in 2 minutes" and wrong-date holds dropped to zero.
Follow-up questions to expect
- "Should the tool raise or return an error?" — Return a structured error for failures the model can react to. Raise only for bugs you want in your error tracker; AutoGen will still pass the text to the model.
- "How do you stop the agent retrying forever?" — Cap
max_tool_iterations, count failures in the tool and return "do not retry", and keep a token or message cap on the team. - "What about a tool that returns 50 of 4,000 rows?" — Say so in the result, with a cursor, so the agent does not summarise 50 rows as if they were all of them.