Course Content
AutoGen Essentials
7 sections · 28 lessons
How do you persist agent state across sessions (conversation logs, summaries, vector memory)?
What you need to know
Three kinds of persisted data
| Data | Store | Used for |
|---|---|---|
| Transcript (every message) | Postgres or object storage, append-only | Audit, debugging, evals; not reloaded into context |
| Agent or team state | Redis or Postgres, keyed by session | Resuming a paused run exactly |
| Summaries and facts | Tables (structured), vector store (open recall) | Starting the next session with what matters |
Save and load state in 0.4+
1import json23async def pause(team, session_id: str) -> None:4 state = await team.save_state() # plain dict: agent contexts + team state5 await redis.set(f"team:{session_id}", json.dumps(state), ex=7 * 24 * 3600)67async def resume(session_id: str, reply: str):8 team = build_support_team() # rebuild agents, tools, model clients9 saved = await redis.get(f"team:{session_id}")10 if saved:11 await team.load_state(json.loads(saved))12 return await team.run(task=reply)save_state()exists on agents and on teams. For a team it includes each participant's state and the manager's state (message thread, current turn).- State does not include model clients, tools or credentials. You rebuild those in code and load state into them.
dump_component()/load_component()save the configuration (which agents, prompts, model settings) as JSON. Keep config in git, state in your store.- Do not save state while the team is running; save after
runreturns.
Legacy 0.2: you stored groupchat.messages and each agent's chat_messages yourself, and could pass summary_method="reflection_with_llm" to initiate_chat to get a summary to carry into the next chat.
Sharp edges
- Version your state. If you rename an agent or change its type, old state may not load. Store a schema version and migrate or discard.
- Expire it. Give saved state a TTL; stale sessions should start fresh.
- Do not replay everything. Loading a 200-message state into a new run is expensive and brings back already-corrected mistakes. For long-lived users, prefer summary plus facts.
A real-life example
A broadband company's customer-support triage team handles "my internet is down" tickets. A technician visit is booked, and the customer returns the next day: "The technician came, still not working."
Version one kept the team in memory, so after a pod restart the history was gone and the agent asked for the account number again. Version two:
- On each pause,
save_state()goes to Redis for 7 days. - A 150-token summary ("line fault ticket 55102, technician visited 24 Sep, router replaced") and structured facts (ticket id, router model) go to Postgres.
- The full transcript goes to object storage for audit.
When the customer returns within a week, the team loads state and continues. After a week, a new team starts with only the summary and facts. Repeat questions dropped from 38% of returning chats to 4%.
Follow-up questions to expect
- "What if the process crashes mid-run?" — AutoGen teams do not checkpoint automatically during a run. Save state after each
run(or eachmax_turnsstep), make tools idempotent, and accept that the current turn may repeat. Microsoft Agent Framework and LangGraph have built-in checkpointing if you need more. - "Why not store the transcript and reload it?" — It is long, costly and includes mistakes that were corrected later. A summary plus facts is smaller and cleaner.
- "How do you handle user data deletion?" — Key every store by user and session, so a deletion request can remove transcripts, state, summaries and memories together.