Course Content
CrewAI Multi-Agents
9 sections · 53 lessons
How do you handle disagreements or conflicting outputs between agents?
What you need to know
Two agents disagree when they read different inputs, read the same input differently, or one of them made an error. You want to know which, not just pick a winner.
Step 1: make each side a structured claim
1class Claim(BaseModel):2 answer: str # e.g. "covered" / "not_covered"3 confidence: float # 0 to 1, the agent's own estimate4 evidence: list[str] # clause numbers, record IDs, URLs5 source_checked: bool # did a tool return this, or is it from memory?A model's own confidence is only a rough hint, so the evidence field matters more.
Step 2: resolve with rules in code
- Agree — both answers match: accept.
- One has checked evidence — prefer the claim whose evidence came from a tool (a database lookup, the policy PDF) over one from the model's memory.
- Evidence can be re-checked — re-fetch the cited record in code and see who is right.
- Real conflict — both have checked evidence that disagrees: escalate to a person with both claims side by side.
These rules are cheap, fast and give the same result every time. They also leave an audit trail.
Step 3: an adjudicator only for judgement calls
Some conflicts are about interpretation, like "is a burst pipe sudden damage?". Then an adjudicator task gets both claims in context and must return which claim it chose, the reason, and the evidence it relied on. An adjudicator without a stated reason cannot be audited.
Why not let them debate?
Free-text debate between agents costs many tokens and often ends in favour of whichever agent writes longer, more confident text. It also has no natural end. If you use a debate round at all, allow one round and then apply the rules above.
A real-life example
In the insurance-claims review crew, the policy agent says a ₹1.2 lakh water-damage claim is covered under clause 4.2(b). The fraud agent's summary says "damage is gradual seepage, likely excluded".
The resolver sees both claims cite clause numbers, but only the policy agent's clause came from the policy-lookup tool. The code re-reads clause 4.2(b) and 7.1 (exclusions) and finds that "gradual seepage" is an exclusion. The evidence truly conflicts, so the claim goes to a human adjuster with both claims attached. That took 30 seconds of the adjuster's time instead of a 15-minute re-review.
Over a month, the team saw 8% of claims disagree. Most were in water-damage claims, because the intake task did not extract when the damage started. After adding a damage_start_date field to intake, disagreement fell to 3%.
Follow-up questions to expect
- "Can CrewAI resolve conflicts on its own?" — A hierarchical manager will pick something, but silently. I prefer explicit rules, so the choice is logged and repeatable.
- "Should the adjudicator use a stronger model?" — Often yes. It runs rarely, and its decision is the final one.
- "What if agents keep disagreeing?" — Look at the task contracts first. Repeated disagreement on the same field usually means that field is ambiguous or one agent lacks an input.