AI Agent Frameworks

CrewAI and Collaborative Agents


A team built an agent to produce weekly market briefings. One agent, one prompt: "Research the topic, write a 600-word briefing, then fact-check it and flag anything uncertain."

It worked. Every briefing came back with a tidy note at the bottom: "Fact-check: all claims verified against sources." Every single one, for eleven weeks.

Then a reader queried a figure. The briefing said a competitor's revenue was 340 million euros; the actual filing said 240 million. The agent had written 340, and then — as fact-checker — read its own sentence, found it consistent with its own memory of what it had written, and passed it.

This is not a prompt bug. It is structural. When the writer and the checker are the same model call chain with the same context, the checker's "evidence" is the writer's output. There is no independent view. Asking one agent to critique its own work is asking someone to proofread their own typing — they read what they meant, not what is there.

Multi-agent systems are not about parallelism or speed. They are about giving a task a second pair of eyes that has not already seen the answer.

One prompt, or three rolesOne agent, one prompt• Research, write and check in one context• The checker approves its own draft• Trade-offs collapse into one voice• Cheap, and quietly self-confirmingThree roles, separate contexts• Each role has a goalthat names a trade-off• The fact-checker never saw the writing• Handoff is explicit and inspectable• Context is lost unless you pass it on
Separation buys an independent critic; the price is every fact the handoff forgets to carry.

What separating agents actually buys

Three things, and it is worth being precise about them because two of them are often claimed and only sometimes delivered.

Claimed benefitReal?Condition
Independent verificationYes, reliablyThe checker must receive the sources, not just the draft
Focused context per roleYesEach agent gets only its own tools and instructions
Speed through parallelismOnly sometimesRequires genuinely independent tasks; sequential crews are slower than one agent
Better quality overallConditionalGains from specialisation must exceed losses at handoffs

That last row is the honest one. A three-agent crew makes at least three model calls where one agent might make one, and every handoff is a place where information is compressed and lost. If your roles are not genuinely distinct, you pay triple for a worse answer.

CrewAI's model: role, goal, backstory, tools

Bash
pip install crewai crewai-tools
Python
import osfrom crewai import Agent, Task, Crew, Process, LLMfrom crewai_tools import SerperDevTool      # needs SERPER_API_KEYllm = LLM(model=os.getenv("CREW_MODEL", "anthropic/claude-sonnet-5"))search = SerperDevTool()researcher = Agent(    role="Financial Data Researcher",    goal=("Find primary-source figures - filings, official releases, "          "regulator databases - and record the URL and date for each. "          "Prefer one verified number over five unverified ones."),    backstory=("You spent nine years in equity research. You were once "               "burned by a secondary source that misquoted a filing, and "               "since then you refuse to report any figure you have not "               "traced to the original document. You write 'NOT FOUND' "               "rather than estimate."),    tools=[search],    llm=llm,    allow_delegation=False,    max_iter=8,)writer = Agent(    role="Briefing Writer",    goal=("Turn supplied research into a 600-word briefing that a busy "          "executive can read in three minutes. Use only figures present "          "in the research. Never introduce a number of your own."),    backstory=("You write for readers who will act on what you say. You "               "have a hard rule: every number in your copy carries its "               "source in brackets. If the research does not contain a "               "figure you need, you write '[figure unavailable]'."),    tools=[],                       # deliberately no search    llm=llm,    allow_delegation=False,)checker = Agent(    role="Fact Checker",    goal=("Verify every numeric claim and every named entity in the draft "          "against the research notes and, where necessary, the live web. "          "Report discrepancies; do not rewrite the draft."),    backstory=("You are adversarial by profession. You assume every figure "               "is wrong until you have seen it in a source. You have no "               "attachment to the draft - you did not write it."),    tools=[search],    llm=llm,    allow_delegation=False,)

Four fields, and each does specific work:

FieldFunctionWeak versionStrong version
roleNames the specialism"Assistant""Financial Data Researcher"
goalStates the objective and its trade-off"Find good information""Prefer one verified number over five unverified ones"
backstorySupplies checkable standards"You are experienced""You write 'NOT FOUND' rather than estimate"
toolsBounds what it can doAll tools for everyoneWriter has none — it cannot invent new facts

The writer having no tools is the least obvious and most valuable decision here. A writer with a search tool will, when the research is thin, go and find something — and now you have unsourced material entering the draft at a stage nobody is checking. Removing the tool makes "[figure unavailable]" the only available behaviour.

Goals that state a trade-off

Every goal above resolves a conflict rather than expressing an aspiration. "Find good information" tells the agent nothing when it must choose between a fast secondary source and a slow primary one. "Prefer one verified number over five unverified ones" answers exactly that question. When you write a goal, ask: what will this agent have to choose between, and have I told it which way to go?

Tasks and the crew

Python
research_task = Task(    description=("Research {company}'s most recent reported quarter. Find: "                 "revenue, year-on-year growth, operating margin, and one "                 "notable strategic development. For each figure record the "                 "source URL and publication date."),    expected_output=("A markdown list. Each line: METRIC | VALUE | SOURCE URL | "                     "DATE. Use 'NOT FOUND' for anything you cannot verify."),    agent=researcher,)write_task = Task(    description=("Write a 600-word executive briefing on {company} using ONLY "                 "the research notes. Structure: headline finding, three "                 "supporting paragraphs, one-line implication."),    expected_output="600-word briefing, every figure followed by [source].",    agent=writer,    context=[research_task],          # receives the researcher's output)check_task = Task(    description=("Check every numeric claim and named entity in the briefing "                 "against the research notes. Where a figure is not in the "                 "notes, search for it. List discrepancies."),    expected_output=("A table: CLAIM | STATUS (verified/wrong/unverifiable) | "                     "EVIDENCE. Then a verdict line: PASS or FAIL."),    agent=checker,    context=[research_task, write_task],   # sees BOTH, not just the draft)crew = Crew(    agents=[researcher, writer, checker],    tasks=[research_task, write_task, check_task],    process=Process.sequential,    verbose=True,)result = crew.kickoff(inputs={"company": "Delivery Hero"})

context=[research_task, write_task] on the checker is the fix for the opening failure. If the checker received only the draft, it would be verifying prose against nothing. Receiving the research notes as well lets it do the one thing that matters: compare the number in the copy against the number in the source.

Three coordination patterns

Sequential handoff

A → B → C, each receiving the previous output. Simple, predictable, and the default. Its weakness is compression: each handoff passes a summary, and detail present at step 1 is gone by step 3.

Quantify it. Suppose the researcher produces 2,000 tokens of notes with 14 figures. The writer's brief is 600 words — about 800 tokens — and includes 6 figures. If the checker sees only the briefing, 8 figures have disappeared and the ones that remain cannot be checked against anything. Passing both artefacts costs 2,800 input tokens instead of 800: an extra 0.006 dollars per briefing at an example rate of 3 dollars per million. For catching a 100-million-euro error, that is the cheapest insurance in the system.

Hierarchical

A manager agent decides who does what and may send work back.

Python
crew = Crew(    agents=[researcher, writer, checker],    tasks=[briefing_task],    process=Process.hierarchical,    manager_llm=llm,    max_rpm=20,)

Powerful and dangerous. The manager can loop: writer submits, checker rejects, writer revises, checker rejects again. Without a bound this runs until your budget does. CrewAI gives you max_iter per agent and max_rpm per crew; use both, and additionally track revision count in your own wrapper:

Python
MAX_REVISIONS = 2revisions = 0while revisions < MAX_REVISIONS:    verdict = run_check(draft)    if verdict.startswith("PASS"):        break    draft = run_revision(draft, verdict)    revisions += 1else:    draft += "\n\n[UNRESOLVED after 2 revisions: " + verdict + "]"

Note the else clause. Failing loudly after two revisions is far better than quietly shipping the third attempt as though it passed.

Parallel then synthesise

Several agents work independently on different aspects; one merges. This is the only pattern that actually saves wall-clock time.

Python
for t in (market_task, competitor_task, regulatory_task):    t.async_execution = Truesynthesis_task.context = [market_task, competitor_task, regulatory_task]

With three research tasks at 22 seconds each and a 15-second synthesis: sequential is 3×22+15=813 \times 22 + 15 = 81 seconds; parallel is max⁡(22,22,22)+15=37\max(22, 22, 22) + 15 = 37 seconds. A 2.2x improvement — but only because the three tasks genuinely do not need each other's output. Mark a dependent task async_execution=True and it will run against missing context and produce confident nonsense.

PatternWall clockModel callsUse whenMain risk
SequentialSum of stepsOne per taskReal dependencies between stagesContext loss at handoffs
HierarchicalVariableTask calls + manager callsWork allocation is not known upfrontUnbounded revision loops
Parallel + synthesiseSlowest branch + mergeOne per branch + mergeBranches are independentDuplicate work; contradictory findings

Designing roles that do not collide

Here is a crew that looks sensible and performs badly:

Text
Agent 1: "Research Analyst"     - researches the topicAgent 2: "Market Analyst"       - researches the marketAgent 3: "Industry Expert"      - provides industry context

All three will search the web for overlapping queries, all three will return overlapping findings, and the synthesiser will receive three versions of the same material with slightly different numbers. You paid triple for redundancy plus a contradiction problem.

The test for whether two roles are genuinely distinct: write down one question each agent can answer that the other cannot. If you cannot, they are one agent.

CollisionSymptomFix
Two researchersNear-identical findings, conflicting figuresSplit by source type: filings vs news vs regulator, not by topic
Writer with search accessUnsourced figures appear in the draftRemove the tool
Checker without sourcesEverything passesPass the research task as context
Everyone has all toolsRoles converge in behaviourMinimum viable toolset per agent

Splitting by source type rather than topic is the practical trick. "Find it in the annual filing" and "find it in the last 30 days of news" are questions that cannot be confused. "Research the company" and "research the market" can.

The context-loss problem, measured

Handoffs lose information, and the loss is compounding. A realistic chain:

StageArtefactFigures presentSources cited
Researcher output2,000 tokens of notes1414
Writer output800-token briefing66
Checker input (draft only)800 tokens66, unverifiable
Checker input (draft + notes)2,800 tokens6 checkable against 14Traceable

Two mitigations are worth building in from the start. First, pass structured artefacts rather than prose between agents — a table of METRIC | VALUE | SOURCE | DATE survives a handoff intact where a paragraph does not. Second, keep a shared store that every agent writes to and reads from, so the artefact chain is not the only channel:

Python
class CrewNotes:    def __init__(self): self.facts = {}    def record(self, metric, value, source, date, by):        prior = self.facts.get(metric)        if prior and prior["value"] != value:            self.facts[metric]["conflict"] = (prior, {"value": value,                                                      "source": source, "by": by})        else:            self.facts[metric] = {"value": value, "source": source,                                  "date": date, "by": by}    def conflicts(self):        return {k: v for k, v in self.facts.items() if "conflict" in v}

The conflicts() method is the point. When two agents record different values for the same metric, that disagreement is the single most valuable signal the crew produces — and in a plain handoff chain it is invisible, because the second value simply overwrites the first.

Give each agent the minimum tools its role requires. A tool is not just a capability - it is a licence to introduce material that no later stage will question.

Four failure modes with names

Overlapping roles

Covered above. Symptom: your crew of four produces the output of one, at four times the cost, with internal contradictions.

The self-approving checker

The opening story. Symptom: a verification step that has never once failed. If your fact-checker has a 100 per cent pass rate over eleven runs, it is not checking anything. Deliberately inject a wrong figure into a draft and confirm the checker catches it — a verification step you have not tested against a known error is a decoration, not a control.

Unbounded revision loops

Manager sends back, writer revises, checker rejects, repeat. Two revisions maximum, then escalate with the unresolved complaint attached. The cost of an unbounded loop is not just tokens: it is that the third and fourth revisions typically drift away from the source material as the writer tries increasingly hard to satisfy a checker's phrasing rather than the facts.

Assignment by rotation rather than fit

Hierarchical crews with a weak manager prompt distribute work round-robin — each agent gets a turn regardless of suitability. Symptom: your researcher writing prose and your writer running searches. Fix it by giving the manager the roles explicitly and requiring a justification:

Python
manager_prompt = (    "Assign each task to exactly one agent. Available agents and the ONLY "    "things they are good at:\n"    "- Financial Data Researcher: locating and verifying primary-source figures\n"    "- Briefing Writer: turning verified notes into executive prose\n"    "- Fact Checker: comparing claims against sources\n"    "For each assignment state, in one clause, why that agent and not the "    "others. Never assign a task to an agent whose listed strength does not "    "cover it - if none fits, say so.")

Two beliefs worth abandoning

"More agents means better output." Each additional agent adds a handoff, and every handoff loses information and costs a model call. The quality curve peaks early — usually at two or three genuinely distinct roles — and declines after that. Before adding a fourth agent, name the question it will answer that no existing agent can.

"The backstory is flavour text." The backstory is the highest-leverage field on the agent, because it is where you put behaviour that is hard to state as a goal. "You write 'NOT FOUND' rather than estimate" changes measurable output: in side-by-side runs, researchers with that sentence report unverifiable metrics as missing, and researchers without it produce plausible numbers with no source. It reads like character-writing. It functions as a constraint.

Deciding whether you need a crew at all

Start with one agent. Add a second only when you can point at a failure that a second, independent view would catch — and the canonical one is verification, because a single agent cannot check its own work. That is not a limitation of the prompt; it is a limitation of having only one view of the evidence.

When you do build a crew, three decisions determine whether it is worth the cost. Give each agent the minimum tools its role requires, because a tool is a licence to invent. Pass structured artefacts between stages, not prose, because structure survives handoffs and prose does not. And cap every loop — max_iter per agent, a revision counter around any manager-worker cycle, and a wall-clock deadline on the crew — because the failure mode of a coordination system is never a crash, it is a polite conversation between three agents that never ends.

Finally, measure the thing that justified the crew in the first place. If you added a fact-checker, track its rejection rate. A rate near zero means the checker is not working; a rate near one means the writer is not. Either is worth knowing, and neither is visible in the finished briefing, which will look equally confident in both cases.