AutoGen Essentials

Course Content

AutoGen Essentials

7 sections · 28 lessons

How do you make an AutoGen agent decide when to call a tool vs respond directly?


What you need to know

How the decision happens

On each call, AutoGen sends the model the conversation plus the JSON schemas of the agent's tools. The model returns either text or one or more tool call requests. There is no separate "router": the only inputs to the decision are your system message, your tool descriptions and the conversation.

What moves the decision

  • Descriptions with boundaries. "Search live flight prices between two airports on a date. Do not use for general travel advice or visa questions."
  • A policy line. "If the answer needs live prices, availability or the user's bookings, call a tool. For general knowledge, answer directly."
  • Fewer tools per agent. Tool choice gets worse as the list grows and names look alike. Five clear tools beat twenty.
  • Plain text is a valid action. Say so. Otherwise models tend to over-call.

The AutoGen settings that bound it

Python
agent = AssistantAgent(    "travel_helper",    model_client=client,    tools=[search_flights, get_weather],    max_tool_iterations=3,        # at most 3 rounds of tool calls per turn    reflect_on_tool_use=True,     # one more model call to explain results    system_message=(        "Answer general travel questions yourself. Call search_flights only "        "when the user gives cities and a date. Never invent prices."    ),)
  • max_tool_iterations (default 1): how many rounds of "call tools, see results" one turn may take. Raise it for chained lookups; keep it low so a confused agent cannot retry forever.
  • reflect_on_tool_use: True gives a natural-language answer after tools run. False returns a ToolCallSummaryMessage built from tool_call_summary_format (default "{result}"), which is cheaper and fine when another agent will read the raw result.

Measure it like a classifier

Build 50 to 100 prompts labelled with the expected tool, including "no tool". Track:

ErrorSymptomUsual fix
Over-callingTool called when not needed; slower, costlierSay plain answers are fine; narrow the description
Under-callingAnswer made up instead of looked upPolicy line "never guess"; clearer "use when"
Wrong toolSimilar tools confusedRename, merge or split tools across agents

A real-life example

A travel app's helper answered "Is Goa good to visit in July?" by calling search_flights with made-up dates, costing 2 seconds and returning nothing useful. Yet for "cheapest flight Bengaluru to Goa on 12 July" it sometimes quoted "around ₹3,500" from memory.

The team labelled 80 real questions: 45 needed no tool, 30 needed search_flights, 5 needed get_weather. Before the change, over-calls were 11 of 45 and under-calls 6 of 30. After rewriting the two descriptions and adding "Never state a price you did not get from search_flights", over-calls fell to 2 and under-calls to 1. They kept max_tool_iterations=3 because some questions need weather and then flights.

Follow-up questions to expect

  • "Can you force a tool call?" — Many model APIs support a tool-choice setting, but in an agent it is usually better to split the step: a dedicated agent or a plain function call when a lookup must always happen.
  • "Why not let the agent call as many tools as it likes?" — Unbounded loops cost money and hide failures. A cap turns a stuck agent into a visible, finished turn.
  • "Does the model call several tools at once?" — It can request parallel calls in one round; AutoGen runs them and returns all results. Turn it off on the model client if order matters.