CrewAI Multi-Agents

Course Content

CrewAI Multi-Agents

9 sections · 53 lessons

How do you retry or reassign failed tasks?


Retry at the lowest layer that can fix itTool — backoff on timeouts and 429sTask — guardrail feedback, then retryAgent — max_retry_limit on errorsRun — resume from checkpoint or replayFlow — router to a stronger crew or a person
A network blip should never rerun a whole crew, and every layer that retries must hit idempotent tools.

What you need to know

Different failures need different retries. Retrying a whole crew because one API call timed out wastes money and time.

LayerWhat failsHow to retry
Tooltimeout, HTTP 429, 5xxbackoff inside the tool's _run
LLMprovider errorLLM(max_retries=...)
Task outputwrong format or contentguardrail / guardrails + guardrail_max_retries (default 3)
Agentexecution errorsmax_retry_limit (default 2)
Crew runcrash, deploy, timeoutcheckpoint=True and resume, or crew.replay(task_id=...)
Whole jobstill failingFlow @router to a stronger model or a human

Guardrail retries are targeted

A guardrail's failure message goes back to the agent, so the retry is a correction ("the amount must be a number in rupees"), not a blind re-roll. This is why guardrail retries often succeed on the second try.

Reassignment with a Flow

Python
from crewai.flow.flow import Flow, start, router, listenclass ClaimReview(Flow[ClaimState]):    @start()    def standard_review(self):        try:            out = standard_crew.kickoff(inputs={"claim": self.state.claim})            self.state.decision = out.pydantic        except Exception as e:            self.state.error = str(e)    @router(standard_review)    def check(self):        return "escalate" if self.state.decision is None else "done"    @listen("escalate")    def senior_review(self):        out = senior_crew.kickoff(inputs={"claim": self.state.claim})        self.state.decision = out.pydantic   # stronger model, stricter tools

senior_crew is the same tasks with a stronger model; a second failure there could route to a human queue.

Side effects and idempotency

Retries run tools again. If a tool sends an SMS or writes a payment row, give each action a key (like claim_id + "approved") and skip it if the key already exists.

A real-life example

An insurance-claims review crew processes 2,000 claims a night. Three kinds of failure show up:

  1. The policy-document API returns 503 for about 2% of calls. The tool now retries 3 times with 1, 2 and 4 second waits. Visible failures drop to 0.1%.
  2. The summary task sometimes omits the deductible. A function guardrail checks the deductible_inr field; about 5% of tasks need one retry, and almost all pass on it.
  3. Complex multi-policy claims fail the standard crew about 1% of the time. The Flow routes them to a senior crew with a stronger model (4 times the cost per claim, but only for 20 claims), and anything that fails there goes to a human adjuster.

Early on, a retried task sent the customer two "claim received" SMS messages. The team added an idempotency key per claim and message type.

Follow-up questions to expect

  • "Why not let the agent handle a 429 itself?" — The model cannot fix a rate limit; it will try something else or invent data. Handle it in code.
  • "What is the difference between replay and checkpoints?" — replay reruns from a task using the last kickoff's saved outputs, which is handy in development. Checkpoints save state during any run and let you resume a specific run after a crash.
  • "How many retries?" — Two or three. If it fails more, the problem is not random, and more retries only cost money.