CrewAI Multi-Agents

Course Content

CrewAI Multi-Agents

9 sections · 53 lessons

How do you implement consensus or voting mechanisms?


Three different models vote on one ticket's priorityP1P1P2012model Amodel Bmodel Cdisagrees2 of 3 means P1 queue with a check flag; 3 of 3 routes automatically.
Voters from different models make different mistakes, so the agreement rate is a real confidence signal rather than an echo.

What you need to know

The key word is independent

Three runs of the same model with the same prompt tend to make the same mistake. Voting helps most when the voters differ: different models, different prompts, or different evidence. In CrewAI, the easy way is three agents with different llm values.

Two ways to build it

Simple Python loop for a batch job:

Python
from collections import Counterfrom statistics import mediandef vote(crews, inputs):    outs = [c.kickoff(inputs=inputs).pydantic for c in crews]    labels = Counter(o.priority for o in outs)    label, count = labels.most_common(1)[0]    return {"priority": label,            "agreement": count / len(outs),            "eta_hours": median(o.eta_hours for o in outs)}

A Flow with parallel branches when latency matters:

Python
from collections import Counterfrom crewai.flow.flow import Flow, start, listen, and_from pydantic import BaseModelclass VoteState(BaseModel):    ticket: str = ""    votes: list[str] = []class PriorityVote(Flow[VoteState]):    def ask(self, crew):        out = crew.kickoff(inputs={"ticket": self.state.ticket})        self.state.votes.append(out.pydantic.priority)    @start()    def voter_a(self): self.ask(crew_a)    @start()    def voter_b(self): self.ask(crew_b)    @start()    def voter_c(self): self.ask(crew_c)    @listen(and_(voter_a, voter_b, voter_c))    def tally(self):        return Counter(self.state.votes).most_common(1)[0]

Several @start() methods begin together, and and_(...) makes tally wait for all three. crew_a, crew_b and crew_c are one-agent crews that differ only in their llm.

How to aggregate

Output typeAggregation
Label (priority, category)majority vote
Number (amount, score)median, which ignores one wild value
Yes/no with high riskrequire unanimity, else escalate
Free textpick the candidate most similar to the others; do not merge

A real-life example

A customer-support escalation crew assigns each escalated ticket a priority: P1 (outage), P2 or P3. A wrong P3 on a real outage means an enterprise customer waits a day.

The team uses three voters: one OpenAI model, one Anthropic model and one open-weight model, each as a small classifier agent. For a ticket saying "all 40 POS terminals in our Bengaluru stores show payment failed since 9 am", the votes are P1, P1, P2. Majority P1, agreement 2 of 3, so it goes to the P1 queue with a "check" flag. Unanimous votes (about 85% of tickets) route automatically. Missed P1 tickets dropped from 6 a month to 1, and the voting costs about ₹0.40 extra per escalated ticket, which is small next to a lost enterprise contract.

Follow-up questions to expect

  • "Isn't this just self-consistency?" — Same idea. Self-consistency samples one model several times; using different models gives more independent errors.
  • "How many voters?" — Three is common: it breaks ties and triples cost. Five rarely helps enough to justify the cost.
  • "Can an LLM do the aggregation?" — For labels and numbers, no, use code. For free text, a judge that picks the most consistent candidate is reasonable.