Course Content
CrewAI Multi-Agents
9 sections · 53 lessons
How do you implement consensus or voting mechanisms?
What you need to know
The key word is independent
Three runs of the same model with the same prompt tend to make the same mistake. Voting helps most when the voters differ: different models, different prompts, or different evidence. In CrewAI, the easy way is three agents with different llm values.
Two ways to build it
Simple Python loop for a batch job:
1from collections import Counter2from statistics import median34def vote(crews, inputs):5 outs = [c.kickoff(inputs=inputs).pydantic for c in crews]6 labels = Counter(o.priority for o in outs)7 label, count = labels.most_common(1)[0]8 return {"priority": label,9 "agreement": count / len(outs),10 "eta_hours": median(o.eta_hours for o in outs)}A Flow with parallel branches when latency matters:
1from collections import Counter2from crewai.flow.flow import Flow, start, listen, and_3from pydantic import BaseModel45class VoteState(BaseModel):6 ticket: str = ""7 votes: list[str] = []89class PriorityVote(Flow[VoteState]):10 def ask(self, crew):11 out = crew.kickoff(inputs={"ticket": self.state.ticket})12 self.state.votes.append(out.pydantic.priority)1314 @start()15 def voter_a(self): self.ask(crew_a)16 @start()17 def voter_b(self): self.ask(crew_b)18 @start()19 def voter_c(self): self.ask(crew_c)2021 @listen(and_(voter_a, voter_b, voter_c))22 def tally(self):23 return Counter(self.state.votes).most_common(1)[0]Several @start() methods begin together, and and_(...) makes tally wait for all three. crew_a, crew_b and crew_c are one-agent crews that differ only in their llm.
How to aggregate
| Output type | Aggregation |
|---|---|
| Label (priority, category) | majority vote |
| Number (amount, score) | median, which ignores one wild value |
| Yes/no with high risk | require unanimity, else escalate |
| Free text | pick the candidate most similar to the others; do not merge |
A real-life example
A customer-support escalation crew assigns each escalated ticket a priority: P1 (outage), P2 or P3. A wrong P3 on a real outage means an enterprise customer waits a day.
The team uses three voters: one OpenAI model, one Anthropic model and one open-weight model, each as a small classifier agent. For a ticket saying "all 40 POS terminals in our Bengaluru stores show payment failed since 9 am", the votes are P1, P1, P2. Majority P1, agreement 2 of 3, so it goes to the P1 queue with a "check" flag. Unanimous votes (about 85% of tickets) route automatically. Missed P1 tickets dropped from 6 a month to 1, and the voting costs about ₹0.40 extra per escalated ticket, which is small next to a lost enterprise contract.
Follow-up questions to expect
- "Isn't this just self-consistency?" — Same idea. Self-consistency samples one model several times; using different models gives more independent errors.
- "How many voters?" — Three is common: it breaks ties and triples cost. Five rarely helps enough to justify the cost.
- "Can an LLM do the aggregation?" — For labels and numbers, no, use code. For free text, a judge that picks the most consistent candidate is reasonable.