Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your AI assistant becomes dramatically worse during long conversations because context windows fill with irrelevant history. How do you manage conversational memory and context compression effectively?
What you need to know
Why long chats get worse
Every turn adds text. By turn 60, the context holds greetings, abandoned ideas and repeated explanations. The model has more to read (slower, costlier) and more distractions, and details from early turns get less attention. Simple summarisation makes it worse in its own way: "the user gave an order number" survives, the number itself does not.
The layout of a curated context
| Part | Content | Changes |
|---|---|---|
| Stable prefix | System prompt, tools, instructions | Never within a session, so prompt caching keeps working |
| Pinned facts | Order ids, names, dates, amounts, stated preferences, promises made | Slots are updated, never summarised |
| Running state | Decisions, constraints, open questions | Rewritten at each compaction |
| Retrieved history | Older turns relevant to the current message | Fetched on demand |
| Recent turns | The last 6–10 turns word for word | Slides forward |
1def build_context(session, message, budget=24000):2 parts = [SYSTEM_PREFIX, # stable, cacheable3 render_facts(session.pinned_facts), # {"order_id": "OD-55821", ...}4 render_state(session.state)]5 if refers_to_past(message): # "like we discussed earlier"6 parts.append(render_turns(session.search_history(message, k=3)))7 parts.append(render_turns(session.turns[-8:]))8 if count_tokens(parts) > 0.6 * budget:9 session.compact() # rewrite state from turns, move old turns to the searchable store10 return parts + [message]Volatile content goes at the end, so the prefix stays identical and cacheable. Compaction happens only when the context passes 60% of the budget; compacting every turn wastes tokens and keeps changing the context.
- Extract pinned facts every turn — a small model fills slots.
- Keep recent turns verbatim — recent wording carries nuance.
- Compact on a threshold — rewrite the state object from the turns being moved out.
- Retrieve old turns only when the message refers back.
- Avoid drift — regenerate the state from original turns now and then, instead of summarising summaries.
Measure tokens per turn, cost per session, and a memory-recall eval: scripted long conversations that ask at turn 60 about a fact stated at turn 5.
A real-life example
Scenario, numbers made up. An airline's support chat averages 45 turns for rebooking cases. Past turn 30, the assistant forgets booking references and repeats questions; satisfaction for long chats is 2.9 out of 5 versus 4.2 for short ones. Cost per long chat is high because the full transcript is re-sent every turn.
The team adds pinned slots (PNR, passenger names, flight dates, fare difference quoted), keeps 8 recent turns, and compacts at 60% of budget. On a 100-conversation recall eval, questions about turn-5 facts asked at turn 50 go from 58% correct to 97%. Tokens per turn fall by 55%, and satisfaction for long chats rises to 3.9.
Follow-up questions to expect
- "Why not just use a million-token context model?" — Cost and latency grow with every turn, and models still use cluttered long context less reliably. Curation helps even when everything would fit.
- "How is this different from long-term memory?" — This is within one session. Long-term memory stores facts across sessions and needs rules for going stale.
- "What if compaction drops something important?" — The original turns stay in storage and can be retrieved; and anything that must never be lost belongs in a pinned slot.