AutoGen Essentials

Course Content

AutoGen Essentials

7 sections · 28 lessons

How do you decide the right number of agents for a workflow (2 vs 3 vs many)?


What you need to know

What each extra agent costs

In a shared group chat every message goes to every agent, and each turn is a model call that re-reads the history. Roughly:

Text
cost per run ≈ turns × (average history length + prompt) × price per token

Adding an agent adds turns and makes the history longer, so cost grows faster than the agent count. Latency grows with turns, because turns are sequential.

A decision ladder

AgentsUse whenExample
1 + toolsThe task is one skill with lookupsAnswer order-status questions
2You need independent checking or real executionWriter + reviewer; analyst + code executor
3A distinct third skill, or a tiebreakPlanner + researcher + writer
Many (5+)Truly different specialists that rarely talk to each otherNested teams, Swarm handoffs or GraphFlow, not one big SelectorGroupChat

Two tests for adding an agent

  • Can you write its prompt without mentioning another agent's job? If not, it is a paragraph in an existing prompt.
  • Does removing it lower the eval score? Run the same golden tasks with and without it. If the score does not drop, it is decoration you are paying for.

Why many agents hurt routing

With SelectorGroupChat, a model reads every participant's description and picks one. With 3 clear roles this is easy. With 9 overlapping roles, the selector often picks the wrong one, and each wrong pick wastes a full turn.

A real-life example

An online electronics store built a customer-support triage team with six agents: greeter, classifier, orders, returns, billing, and summariser. Median cost was ₹4.10 per ticket and median time 38 seconds.

They ran an ablation on 200 past tickets:

  • Removing greeter changed nothing. Removed.
  • classifier and the selector did the same job. They replaced both with a selector_func that routes by the ticket's category field, falling back to the LLM selector.
  • summariser only mattered for tickets escalated to humans; it became a function called at escalation.

The final team was three specialists (orders, returns, billing) behind a code-based router. Resolution accuracy stayed at 91%, median cost fell to ₹1.60, and median time to 14 seconds.

Follow-up questions to expect

  • "When is one agent better than two?" — When there is nothing independent to check and no code to run: for example, a FAQ bot with a search tool. A second agent then only adds cost.
  • "How do you run many specialists without one huge chat?" — Put each group in a nested team or use Swarm handoffs, so each specialist only sees what it needs and returns one result.
  • "Does a bigger model change the answer?" — Often yes. A stronger model with good tools can replace two or three weaker specialised agents, so re-test the team size when you upgrade models.