Course Content
CrewAI Multi-Agents
9 sections · 53 lessons
When should you choose CrewAI over a single-agent architecture?
What you need to know
Why a single agent is the default
One agent is cheaper, faster and easier to debug. Every handoff between agents loses some information, because the next agent only sees the text it is given. So the question is not "can I split this?" but "what does splitting buy me?".
Four good reasons to split
- Different skills. One prompt trying to be a researcher, an analyst and a copywriter shows role bleed — for example, research notes written in marketing language.
- Least privilege. Only one agent should hold the database credential or the email-sending tool. Separate agents make that boundary real.
- Independent review. A critic that did not write the draft is less likely to excuse its mistakes.
- Parallel work. Independent sub-tasks (three competitor profiles) can run at the same time.
Reasons that sound good but are not
| Claim | Better tool |
|---|---|
| "Formatting is a separate concern" | output_pydantic on the task |
| "It feels more modular" | plain Python functions |
| "The steps always run in the same order" | a Flow or a script calling one agent per step |
| "The prompt is long" | trim the prompt first; long is not the same as confused |
How to decide with data
- Build a small golden set — 20 to 50 real inputs with known good answers.
- Run the single agent and record where it fails.
- Split only the part that fails, then run the golden set again.
- Keep the split only if quality rises more than cost.
A real-life example
An NBFC reviews loan documents: salary slips, bank statements and ID proofs. The first build was one agent with OCR, a bank-statement parser and a policy document. On 40 test files it got 31 right. The failures had a pattern: when a bank statement showed large cash deposits, the agent both found them and explained them away in the same breath.
The team split off a risk reviewer agent that only sees the extracted figures and the lending policy, and has one job: list anything that breaks the policy. Accuracy rose to 37 of 40. Cost per file went from about Rs 4 to about Rs 7, and time from 25 to 45 seconds. For a loan decision, that trade-off was easy.
They did not split out a "formatting agent" — a Pydantic model on the final task did that job for free.
Follow-up questions to expect
- "How many agents is too many?" — When you cannot explain each agent's unique job in one sentence, or when handoffs lose facts. Most production crews have two to five agents.
- "Isn't a reviewer agent just the same model checking itself?" — It is the same model, but with a different prompt and without the drafting history, which reduces anchoring. Using a different model for the reviewer helps more.
- "What would you measure?" — Task accuracy on the golden set, cost per run, latency, and the number of tool calls per run.