Course Content
CrewAI Multi-Agents
9 sections · 53 lessons
How do you retry or reassign failed tasks?
What you need to know
Different failures need different retries. Retrying a whole crew because one API call timed out wastes money and time.
| Layer | What fails | How to retry |
|---|---|---|
| Tool | timeout, HTTP 429, 5xx | backoff inside the tool's _run |
| LLM | provider error | LLM(max_retries=...) |
| Task output | wrong format or content | guardrail / guardrails + guardrail_max_retries (default 3) |
| Agent | execution errors | max_retry_limit (default 2) |
| Crew run | crash, deploy, timeout | checkpoint=True and resume, or crew.replay(task_id=...) |
| Whole job | still failing | Flow @router to a stronger model or a human |
Guardrail retries are targeted
A guardrail's failure message goes back to the agent, so the retry is a correction ("the amount must be a number in rupees"), not a blind re-roll. This is why guardrail retries often succeed on the second try.
Reassignment with a Flow
1from crewai.flow.flow import Flow, start, router, listen23class ClaimReview(Flow[ClaimState]):4 @start()5 def standard_review(self):6 try:7 out = standard_crew.kickoff(inputs={"claim": self.state.claim})8 self.state.decision = out.pydantic9 except Exception as e:10 self.state.error = str(e)1112 @router(standard_review)13 def check(self):14 return "escalate" if self.state.decision is None else "done"1516 @listen("escalate")17 def senior_review(self):18 out = senior_crew.kickoff(inputs={"claim": self.state.claim})19 self.state.decision = out.pydantic # stronger model, stricter toolssenior_crew is the same tasks with a stronger model; a second failure there could route to a human queue.
Side effects and idempotency
Retries run tools again. If a tool sends an SMS or writes a payment row, give each action a key (like claim_id + "approved") and skip it if the key already exists.
A real-life example
An insurance-claims review crew processes 2,000 claims a night. Three kinds of failure show up:
- The policy-document API returns 503 for about 2% of calls. The tool now retries 3 times with 1, 2 and 4 second waits. Visible failures drop to 0.1%.
- The summary task sometimes omits the deductible. A function guardrail checks the
deductible_inrfield; about 5% of tasks need one retry, and almost all pass on it. - Complex multi-policy claims fail the standard crew about 1% of the time. The Flow routes them to a senior crew with a stronger model (4 times the cost per claim, but only for 20 claims), and anything that fails there goes to a human adjuster.
Early on, a retried task sent the customer two "claim received" SMS messages. The team added an idempotency key per claim and message type.
Follow-up questions to expect
- "Why not let the agent handle a 429 itself?" — The model cannot fix a rate limit; it will try something else or invent data. Handle it in code.
- "What is the difference between replay and checkpoints?" —
replayreruns from a task using the last kickoff's saved outputs, which is handy in development. Checkpoints save state during any run and let you resume a specific run after a crash. - "How many retries?" — Two or three. If it fails more, the problem is not random, and more retries only cost money.