Course Content
AutoGen Essentials
7 sections · 28 lessons
How do you decide the right number of agents for a workflow (2 vs 3 vs many)?
What you need to know
What each extra agent costs
In a shared group chat every message goes to every agent, and each turn is a model call that re-reads the history. Roughly:
cost per run ≈ turns × (average history length + prompt) × price per tokenAdding an agent adds turns and makes the history longer, so cost grows faster than the agent count. Latency grows with turns, because turns are sequential.
A decision ladder
| Agents | Use when | Example |
|---|---|---|
| 1 + tools | The task is one skill with lookups | Answer order-status questions |
| 2 | You need independent checking or real execution | Writer + reviewer; analyst + code executor |
| 3 | A distinct third skill, or a tiebreak | Planner + researcher + writer |
| Many (5+) | Truly different specialists that rarely talk to each other | Nested teams, Swarm handoffs or GraphFlow, not one big SelectorGroupChat |
Two tests for adding an agent
- Can you write its prompt without mentioning another agent's job? If not, it is a paragraph in an existing prompt.
- Does removing it lower the eval score? Run the same golden tasks with and without it. If the score does not drop, it is decoration you are paying for.
Why many agents hurt routing
With SelectorGroupChat, a model reads every participant's description and picks one. With 3 clear roles this is easy. With 9 overlapping roles, the selector often picks the wrong one, and each wrong pick wastes a full turn.
A real-life example
An online electronics store built a customer-support triage team with six agents: greeter, classifier, orders, returns, billing, and summariser. Median cost was ₹4.10 per ticket and median time 38 seconds.
They ran an ablation on 200 past tickets:
- Removing
greeterchanged nothing. Removed. classifierand the selector did the same job. They replaced both with aselector_functhat routes by the ticket's category field, falling back to the LLM selector.summariseronly mattered for tickets escalated to humans; it became a function called at escalation.
The final team was three specialists (orders, returns, billing) behind a code-based router. Resolution accuracy stayed at 91%, median cost fell to ₹1.60, and median time to 14 seconds.
Follow-up questions to expect
- "When is one agent better than two?" — When there is nothing independent to check and no code to run: for example, a FAQ bot with a search tool. A second agent then only adds cost.
- "How do you run many specialists without one huge chat?" — Put each group in a nested team or use
Swarmhandoffs, so each specialist only sees what it needs and returns one result. - "Does a bigger model change the answer?" — Often yes. A stronger model with good tools can replace two or three weaker specialised agents, so re-test the team size when you upgrade models.