Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your AI agent remembers previous conversations — but after weeks, users notice it recalling outdated preferences and irrelevant context. How do you design long-term memory systems that stay useful without becoming noisy or stale?
What you need to know
Why the naive design goes stale
The common first version embeds every past message and retrieves the most similar ones. After a few weeks the store holds both "I prefer a window seat" (March) and "Aisle seats from now on, please" (May). Both are equally "true" to the system, and retrieval returns whichever happens to embed closer to today's question. Old facts never leave, so noise grows with every conversation.
Store facts, not transcripts
A small, cheap model extracts typed items from each conversation:
1from dataclasses import dataclass, field2from datetime import datetime, timedelta, timezone34@dataclass5class Memory:6 user_id: str7 attribute: str # "seat_preference"8 value: str # "aisle"9 kind: str # preference | constraint | entity | decision10 confidence: float11 source_turn: str12 created_at: datetime = field(default_factory=lambda: datetime.now(timezone.utc))13 superseded_by: str | None = None14 expires_at: datetime | None = None1516TTL = {"preference": None, "constraint": None, "decision": timedelta(days=90), "entity": timedelta(days=30)}1718def write(store, m: Memory):19 for old in store.active(m.user_id, m.attribute): # same user + attribute20 store.mark_superseded(old.id, by=m.id)21 if TTL[m.kind]:22 m.expires_at = m.created_at + TTL[m.kind]23 store.insert(m)Prose memories cannot be compared; typed facts can. The key (user_id, attribute) is what makes "aisle" replace "window" automatically.
Decay by type
| Kind of fact | Example | Lifetime |
|---|---|---|
| Stable constraint | "I'm vegetarian", "I work at Infosys" | Keep until contradicted |
| Preference | "Short answers, please" | Keep, but let new signals override |
| Current context | "Working on the Q3 launch" | 30–90 days, then ask to reconfirm |
| One-off detail | "My flight is on Friday" | Days |
The read path
At answer time, score each candidate memory against the current turn:
score = relevance x recency_weight x confidence x importanceInject only the top 5–10. Dumping the whole profile into every prompt is exactly what makes answers feel noisy and off-topic.
- Extract typed facts after each conversation.
- Supersede older facts with the same key on write.
- Expire volatile facts by type.
- Retrieve the few best-scored memories per turn.
- Consolidate nightly: merge duplicates, summarise clusters, drop memories never retrieved in months.
- Show and edit — a "What I remember about you" page.
The last step is the cheapest way to stay correct, and data-protection laws such as GDPR and India's DPDP Act usually require that users can see and delete stored personal data anyway.
A real-life example
Scenario, numbers made up. A travel-booking assistant has stored raw chat snippets for eight months. Users complain that it keeps suggesting window seats and non-veg meals they stopped choosing long ago. The team audits 500 retrieved memories and finds 22% are outdated and another 18% irrelevant to the question asked.
They switch to typed facts with supersession and TTLs, and retrieve the top 6 instead of the top 20. A month later the same audit shows 4% outdated. An A/B test with memory on versus off shows booking completion up 3 points with memory on — proof that memory is earning its cost.
Follow-up questions to expect
- "How do you detect a contradiction when facts are free text?" — Normalise to a fixed set of attributes where you can. For the rest, retrieve similar existing memories on write and ask a small model whether the new one replaces, adds to, or conflicts with them.
- "How do you measure memory quality?" — Memory precision (a human judges whether retrieved items were relevant), stale-fact rate on an audit sample, and task success with memory on versus off.
- "What about a user who says 'forget that'?" — Delete the item and its embedding for real, not just hide it, and log the deletion.