Course Content
AutoGen Essentials
7 sections · 28 lessons
How do you manage short-term vs long-term memory in AutoGen conversations?
What you need to know
Short-term memory: the model context
By default an AssistantAgent uses UnboundedChatCompletionContext: every message stays and is re-sent on every call. In a long run that means cost and latency keep climbing. Pick a bounded context instead:
| Context class | Keeps | Good for |
|---|---|---|
BufferedChatCompletionContext(buffer_size=10) | Last N messages | Chatty agents where only recent turns matter |
HeadAndTailChatCompletionContext(head_size=2, tail_size=8) | First few and last few messages | Keeping the original task plus current state |
TokenLimitedChatCompletionContext(model_client, token_limit=...) | As much recent history as fits a token budget (experimental) | Variable-length messages |
In a group chat each agent keeps its own context, filled with the messages the team broadcast to it, so you can bound each agent separately.
Legacy 0.2 had no model-context classes; you used the transform_messages capability with MessageHistoryLimiter or MessageTokenLimiter, or summarised between chats.
Long-term memory: the Memory protocol
1from autogen_agentchat.agents import AssistantAgent2from autogen_core.memory import ListMemory, MemoryContent, MemoryMimeType3from autogen_core.model_context import HeadAndTailChatCompletionContext45prefs = ListMemory(name="traveller_prefs")6await prefs.add(MemoryContent(content="Vegetarian meals only",7 mime_type=MemoryMimeType.TEXT))8await prefs.add(MemoryContent(content="Prefers aisle seats; avoids red-eye flights",9 mime_type=MemoryMimeType.TEXT))1011planner = AssistantAgent(12 "planner", model_client=client,13 memory=[prefs],14 model_context=HeadAndTailChatCompletionContext(head_size=2, tail_size=8),15)Before each model call, the agent calls each memory's update_context, which queries the store and adds the results to the model context; the run emits a MemoryQueryEvent so you can see what was retrieved. ListMemory adds all its items, in order. ChromaDBVectorMemory (in autogen_ext.memory.chromadb) retrieves the top k items above a score_threshold; autogen-ext also has Redis and mem0 memory. You can write your own class with add, query, update_context and clear.
What belongs where
| Short-term (context) | Long-term (memory store) |
|---|---|
| The current task and recent turns | Stable preferences and profile facts |
| Latest tool results | Decisions made in earlier sessions |
| Bounded and lossy by design | Selective, with source and timestamp |
A real-life example
A travel-planning assistant for a corporate travel desk plans trips over long chats. After 40 turns, each call was sending 22,000 tokens, replies took 9 seconds, and the planner sometimes forgot the original "return by Friday" rule buried in the middle.
The team made two changes:
HeadAndTailChatCompletionContext(head_size=2, tail_size=10): the head keeps the traveller's request and constraints, the tail keeps the latest options.- A
ListMemoryper employee with 4 to 6 facts ("vegetarian", "aisle seat", "company cap ₹6,000 per night"), loaded from their profile table at the start of each session.
Tokens per call dropped to about 6,000, replies to 3 seconds, and "forgot the constraint" complaints stopped, because the constraint now always sits in the head.
Follow-up questions to expect
- "Why not just use a model with a 1-million-token window?" — You still pay for every token on every call, latency rises, and models recall facts in the middle of long contexts less reliably. Bounded context is cheaper and often more accurate.
- "Is vector memory always better than a list?" — No. For a handful of known facts,
ListMemoryor a normal database table is simpler and exact. Use vector retrieval when there are too many memories to include them all. - "Where did 0.2 keep memory?" — In each agent's
chat_messages; long-term memory came from contrib add-ons such as the Teachability capability or your own retrieval code.