AutoGen Essentials

Course Content

AutoGen Essentials

7 sections · 28 lessons

How do you persist agent state across sessions (conversation logs, summaries, vector memory)?


What you need to know

Three kinds of persisted data

DataStoreUsed for
Transcript (every message)Postgres or object storage, append-onlyAudit, debugging, evals; not reloaded into context
Agent or team stateRedis or Postgres, keyed by sessionResuming a paused run exactly
Summaries and factsTables (structured), vector store (open recall)Starting the next session with what matters

Save and load state in 0.4+

Python
import jsonasync def pause(team, session_id: str) -> None:    state = await team.save_state()        # plain dict: agent contexts + team state    await redis.set(f"team:{session_id}", json.dumps(state), ex=7 * 24 * 3600)async def resume(session_id: str, reply: str):    team = build_support_team()            # rebuild agents, tools, model clients    saved = await redis.get(f"team:{session_id}")    if saved:        await team.load_state(json.loads(saved))    return await team.run(task=reply)
  • save_state() exists on agents and on teams. For a team it includes each participant's state and the manager's state (message thread, current turn).
  • State does not include model clients, tools or credentials. You rebuild those in code and load state into them.
  • dump_component() / load_component() save the configuration (which agents, prompts, model settings) as JSON. Keep config in git, state in your store.
  • Do not save state while the team is running; save after run returns.

Legacy 0.2: you stored groupchat.messages and each agent's chat_messages yourself, and could pass summary_method="reflection_with_llm" to initiate_chat to get a summary to carry into the next chat.

Sharp edges

  • Version your state. If you rename an agent or change its type, old state may not load. Store a schema version and migrate or discard.
  • Expire it. Give saved state a TTL; stale sessions should start fresh.
  • Do not replay everything. Loading a 200-message state into a new run is expensive and brings back already-corrected mistakes. For long-lived users, prefer summary plus facts.

A real-life example

A broadband company's customer-support triage team handles "my internet is down" tickets. A technician visit is booked, and the customer returns the next day: "The technician came, still not working."

Version one kept the team in memory, so after a pod restart the history was gone and the agent asked for the account number again. Version two:

  • On each pause, save_state() goes to Redis for 7 days.
  • A 150-token summary ("line fault ticket 55102, technician visited 24 Sep, router replaced") and structured facts (ticket id, router model) go to Postgres.
  • The full transcript goes to object storage for audit.

When the customer returns within a week, the team loads state and continues. After a week, a new team starts with only the summary and facts. Repeat questions dropped from 38% of returning chats to 4%.

Follow-up questions to expect

  • "What if the process crashes mid-run?" — AutoGen teams do not checkpoint automatically during a run. Save state after each run (or each max_turns step), make tools idempotent, and accept that the current turn may repeat. Microsoft Agent Framework and LangGraph have built-in checkpointing if you need more.
  • "Why not store the transcript and reload it?" — It is long, costly and includes mistakes that were corrected later. A summary plus facts is smaller and cleaner.
  • "How do you handle user data deletion?" — Key every store by user and session, so a deletion request can remove transcripts, state, summaries and memories together.