AutoGen Essentials

Course Content

AutoGen Essentials

7 sections · 28 lessons

How do you manage short-term vs long-term memory in AutoGen conversations?


What you need to know

Short-term memory: the model context

By default an AssistantAgent uses UnboundedChatCompletionContext: every message stays and is re-sent on every call. In a long run that means cost and latency keep climbing. Pick a bounded context instead:

Context classKeepsGood for
BufferedChatCompletionContext(buffer_size=10)Last N messagesChatty agents where only recent turns matter
HeadAndTailChatCompletionContext(head_size=2, tail_size=8)First few and last few messagesKeeping the original task plus current state
TokenLimitedChatCompletionContext(model_client, token_limit=...)As much recent history as fits a token budget (experimental)Variable-length messages

In a group chat each agent keeps its own context, filled with the messages the team broadcast to it, so you can bound each agent separately.

Legacy 0.2 had no model-context classes; you used the transform_messages capability with MessageHistoryLimiter or MessageTokenLimiter, or summarised between chats.

Long-term memory: the Memory protocol

Python
from autogen_agentchat.agents import AssistantAgentfrom autogen_core.memory import ListMemory, MemoryContent, MemoryMimeTypefrom autogen_core.model_context import HeadAndTailChatCompletionContextprefs = ListMemory(name="traveller_prefs")await prefs.add(MemoryContent(content="Vegetarian meals only",                              mime_type=MemoryMimeType.TEXT))await prefs.add(MemoryContent(content="Prefers aisle seats; avoids red-eye flights",                              mime_type=MemoryMimeType.TEXT))planner = AssistantAgent(    "planner", model_client=client,    memory=[prefs],    model_context=HeadAndTailChatCompletionContext(head_size=2, tail_size=8),)

Before each model call, the agent calls each memory's update_context, which queries the store and adds the results to the model context; the run emits a MemoryQueryEvent so you can see what was retrieved. ListMemory adds all its items, in order. ChromaDBVectorMemory (in autogen_ext.memory.chromadb) retrieves the top k items above a score_threshold; autogen-ext also has Redis and mem0 memory. You can write your own class with add, query, update_context and clear.

What belongs where

Short-term (context)Long-term (memory store)
The current task and recent turnsStable preferences and profile facts
Latest tool resultsDecisions made in earlier sessions
Bounded and lossy by designSelective, with source and timestamp

A real-life example

A travel-planning assistant for a corporate travel desk plans trips over long chats. After 40 turns, each call was sending 22,000 tokens, replies took 9 seconds, and the planner sometimes forgot the original "return by Friday" rule buried in the middle.

The team made two changes:

  • HeadAndTailChatCompletionContext(head_size=2, tail_size=10): the head keeps the traveller's request and constraints, the tail keeps the latest options.
  • A ListMemory per employee with 4 to 6 facts ("vegetarian", "aisle seat", "company cap ₹6,000 per night"), loaded from their profile table at the start of each session.

Tokens per call dropped to about 6,000, replies to 3 seconds, and "forgot the constraint" complaints stopped, because the constraint now always sits in the head.

Follow-up questions to expect

  • "Why not just use a model with a 1-million-token window?" — You still pay for every token on every call, latency rises, and models recall facts in the middle of long contexts less reliably. Bounded context is cheaper and often more accurate.
  • "Is vector memory always better than a list?" — No. For a handful of known facts, ListMemory or a normal database table is simpler and exact. Use vector retrieval when there are too many memories to include them all.
  • "Where did 0.2 keep memory?" — In each agent's chat_messages; long-term memory came from contrib add-ons such as the Teachability capability or your own retrieval code.