Course Content
AutoGen Essentials
7 sections · 28 lessons
What are effective strategies to keep context windows small while maintaining continuity?
What you need to know
Where tokens go in a multi-agent run
In an AutoGen group chat, every chat message is broadcast to every participant, and each agent re-sends its whole context on each call. So one 3,000-token tool result, seen by 4 agents over 10 turns, can be paid for dozens of times.
Five techniques
- Bound the window —
HeadAndTailChatCompletionContext(head keeps the task and rules, tail keeps current state), orBufferedChatCompletionContextfor simple chats. - Roll a summary — every N turns, write a fixed-size note ("decisions, open questions, current state") and drop the raw turns it covers.
- Stop the broadcast — a nested team,
SocietyOfMindAgentorAgentToolruns a sub-conversation and returns one message instead of twenty. - Keep bulk out — write large results to the work directory or a store; return a path or id and a short summary. With
reflect_on_tool_use=False,tool_call_summary_formatcontrols exactly what lands in the conversation. - Keep the prefix stable — keep the system prompt and tool schemas byte-identical across calls so provider prompt caching applies.
Nesting a team as one participant
1from autogen_agentchat.agents import SocietyOfMindAgent2from autogen_agentchat.teams import RoundRobinGroupChat34inner = RoundRobinGroupChat([analyst, runner],5 termination_condition=done_or_12_messages)6analysis = SocietyOfMindAgent("analysis", team=inner, model_client=client)78outer = RoundRobinGroupChat([planner, analysis, writer],9 termination_condition=final_or_20_messages)The analyst and runner may exchange ten messages with code and tracebacks inside inner. SocietyOfMindAgent then uses its model to write one summary message, and only that goes to planner and writer.
Legacy 0.2 equivalents
register_nested_chats hid a sub-conversation behind one agent; summary_method="reflection_with_llm" summarised a finished chat; transform_messages trimmed history.
A real-life example
A data-analysis agent at an insurance company answers questions over claim data. A typical run: plan, write SQL, run it, fix errors, write pandas, make a chart, explain. Runs averaged 31 messages and 410,000 total tokens, because every traceback and 200-row preview went to the planner and the writer too.
Changes:
- The analyst and code runner moved into an inner team behind a
SocietyOfMindAgent; the outer team saw one summary per analysis step. - Query results over 20 rows were saved as CSV in the sandbox work directory; the tool returned the path plus the first 5 rows and column stats.
- The planner used
HeadAndTailChatCompletionContext(head_size=2, tail_size=6).
Total tokens per run fell to about 95,000. Answer quality on their 40-question eval rose from 31 to 34 correct, because the writer was no longer confused by stale tracebacks.
Follow-up questions to expect
- "Doesn't summarising lose information?" — Yes, on purpose. Keep what later steps need (decisions, ids, constraints) and store the raw text elsewhere in case you must look it up.
- "How does prompt caching relate?" — Providers can cache a repeated prompt prefix, making it cheaper and faster. It only works if that prefix does not change, so do not put timestamps or changing summaries at the top.
- "Head-and-tail or summary?" — Use both: head-and-tail as a hard bound, summary for continuity across the dropped middle.