Agentic AI Patterns

Course Content

Agentic AI Patterns

9 sections · 50 lessons

How does memory work in AI agents (short-term, long-term, episodic)?


Four kinds of agent memorythis task's turnscontext windowpast runs,outcomesDB with timestampsstable factsSQL rowshow-to recipesprompt libraryHoldsLives inShort-termEpisodicSemanticProceduralTwo past Redis incidents pointed the triage agent at eviction rate first.
The store is the easy part; what to write, how to fetch it and when to expire it decide whether memory helps.

What you need to know

The memory types

TypeWhat it holdsWhere it livesExample
Short-term (working)This task's turns, plan, tool resultsContext window, plus a state object"Flights found so far: 3 options"
EpisodicPast events and their outcomesDatabase or vector store, with timestamps"On 12 March, refund was approved for order 881"
SemanticStable facts about users and entitiesStructured database rows"Priya prefers aisle seats"
ProceduralHow to do a task wellPrompt library, validated examples"For Air India changes, call get_fare_rules first"

The three policies

  • Write policy. A step that decides what is worth keeping, usually an extraction call at the end of a task. Writing every message is the most common mistake; it fills memory with noise.
  • Retrieval policy. Fetch by relevance and recency, always filtered by user and tenant, with a hard cap on how many memories enter the prompt (for example, 5).
  • Maintenance. Remove duplicates, expire old items with a TTL, and resolve conflicts. If a new fact contradicts an old one, the newer one with a source wins, and the old one is marked superseded.

Risks

  • Bloat: memories crowd out the actual task.
  • Stale facts stated confidently, like an old address.
  • Poisoning: an injected instruction saved as a "preference" and replayed forever.
  • Cross-user leakage through a missing filter.

Controls: namespace every memory by user and tenant, store the source and timestamp with each one, set TTLs, and let users view and delete what is stored about them.

A real-life example

A DevOps incident-triage agent at a food-delivery company:

  • Short-term: during an incident, the alert, the last 30 minutes of metrics, and the tool results so far.
  • Episodic: after each incident closes, a memory job writes one record: symptoms, root cause, fix, and time to resolve. After six months there are 240 records.
  • Semantic: service ownership ("orders-api is owned by the checkout team"), and known fragile dependencies.
  • Procedural: a few validated triage sequences, such as "for Redis latency, check eviction rate before network".

When orders-api latency spikes on a Friday evening, the agent retrieves the three most similar past incidents. Two were caused by the Redis cache hitting its memory limit during dinner peaks. It checks eviction rate first, finds it at 35,000 per second, and proposes the same fix as last time, with links to both old incidents. Time to first hypothesis: 2 minutes instead of 15.

Follow-up questions to expect

  • "Is a longer context window a replacement for memory?" — No. It is still per-request, costs money each call, and models use very long contexts unevenly. Memory selects what matters.
  • "How do you stop memory poisoning?" — Write memories only from trusted sources or through a validated extraction step, keep provenance, and never store raw instructions from documents as preferences.
  • "How do you handle a user asking to be forgotten?" — Store memories keyed by user so they can be deleted completely, including any vector entries and caches.