Course Content
Agents & Tools Interview Prep
6 sections · 40 lessons
What types of memory do AI agents use, and how are they different?
What you need to know
By lifetime
Short-term (working)
- The current context window
- Messages, tool calls, tool results
- Exact, but limited in size
- Gone when the session ends
Long-term
- Stored in a database, files or a vector store
- Survives across sessions
- Must be retrieved back into context
- Can grow stale or wrong
By content
| Type | Holds | Example | Typical storage |
|---|---|---|---|
| Episodic | Past events | "On 3 Aug the user rebooked a cancelled IndiGo flight" | Logs, summaries, vector search |
| Semantic | Stable facts | "Home airport: BLR; vegetarian meals" | Small structured profile (key-value) |
| Procedural | How to act | "For refunds, check fare rules first" | System prompt, skills, playbooks |
Memory through tools
Long-term memory is usually exposed as tools, so the model decides when to read or write it:
recall(query)/remember(fact)backed by your database.- Anthropic's memory tool (
{"type": "memory_20250818", "name": "memory"}) gives the model file-style commands — view, create, edit — on a/memoriesdirectory. It is a client-side tool: your code implements the storage, so you decide where memories live and who can read them.
Treat memory writes like any side-effect tool: validate them, scope them to the user, and log them.
Design rules
- Keep a small profile always loaded (a few hundred tokens of semantic facts) and retrieve episodic details only when relevant.
- Store facts, not transcripts. "Prefers aisle seat" is useful; a 40-message chat log is noise.
- Consent and control. Tell users what is remembered and let them delete it.
- Never store secrets — card numbers, OTPs, passwords — in memory.
- Expiry (TTL). Old facts go stale: an address from 2023 may be wrong now.
- Conflict handling. Newer facts should replace older ones, not sit beside them.
A real-life example
A travel-booking assistant stores three kinds of memory for each user:
- Semantic profile: home airport BLR, vegetarian, aisle seat, corporate travel policy "economy under 4 hours". About 150 tokens, loaded on every conversation.
- Episodic: trip summaries — "Dec 2025: Goa, stayed in Candolim, disliked the noise". Retrieved when the user asks about a similar trip.
- Procedural: the playbook for rebooking cancelled flights, in the system prompt.
When the user says "Book something like last Goa trip but quieter", the agent calls recall("Goa trip"), finds the note about noise, and searches hotels in South Goa instead.
One incident shaped their rules: a user typed their card number in chat, and an early version saved "card ending 4412, full number …" as a fact. They added a filter that blocks card-like and OTP-like patterns from memory writes, and a settings page listing everything stored.
Follow-up questions to expect
- "Isn't a 1-million-token context enough memory?" — It holds one session, costs money on every turn, and is gone next session. Long-term memory is about what to carry forward, not how much fits.
- "Vector store or key-value store?" — Key-value or a small table for stable facts; vector search for fuzzy recall over many past episodes.
- "How do you test memory?" — Scripted multi-session tests: set a fact in session 1, change it in session 2, and check session 3 uses the newest value.