Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your AI assistant becomes dramatically worse during long conversations because context windows fill with irrelevant history. How do you manage conversational memory and context compression effectively?


A curated context, bottom to topStable prefix:system and toolsPinned facts:PNR, names, datesRunningstate, rewrittenOlder turns,only if referencedLast 8 turns,word for wordtopbottomCompaction runs at 60 percent of the budget, and the original turns stay searchable.
Exact values live in slots a summary cannot touch, and the unchanging bottom layer keeps prompt caching paying.

What you need to know

Why long chats get worse

Every turn adds text. By turn 60, the context holds greetings, abandoned ideas and repeated explanations. The model has more to read (slower, costlier) and more distractions, and details from early turns get less attention. Simple summarisation makes it worse in its own way: "the user gave an order number" survives, the number itself does not.

The layout of a curated context

PartContentChanges
Stable prefixSystem prompt, tools, instructionsNever within a session, so prompt caching keeps working
Pinned factsOrder ids, names, dates, amounts, stated preferences, promises madeSlots are updated, never summarised
Running stateDecisions, constraints, open questionsRewritten at each compaction
Retrieved historyOlder turns relevant to the current messageFetched on demand
Recent turnsThe last 6–10 turns word for wordSlides forward
Python
def build_context(session, message, budget=24000):    parts = [SYSTEM_PREFIX,                                    # stable, cacheable             render_facts(session.pinned_facts),               # {"order_id": "OD-55821", ...}             render_state(session.state)]    if refers_to_past(message):                               # "like we discussed earlier"        parts.append(render_turns(session.search_history(message, k=3)))    parts.append(render_turns(session.turns[-8:]))    if count_tokens(parts) > 0.6 * budget:        session.compact()     # rewrite state from turns, move old turns to the searchable store    return parts + [message]

Volatile content goes at the end, so the prefix stays identical and cacheable. Compaction happens only when the context passes 60% of the budget; compacting every turn wastes tokens and keeps changing the context.

  1. Extract pinned facts every turn — a small model fills slots.
  2. Keep recent turns verbatim — recent wording carries nuance.
  3. Compact on a threshold — rewrite the state object from the turns being moved out.
  4. Retrieve old turns only when the message refers back.
  5. Avoid drift — regenerate the state from original turns now and then, instead of summarising summaries.

Measure tokens per turn, cost per session, and a memory-recall eval: scripted long conversations that ask at turn 60 about a fact stated at turn 5.

A real-life example

Scenario, numbers made up. An airline's support chat averages 45 turns for rebooking cases. Past turn 30, the assistant forgets booking references and repeats questions; satisfaction for long chats is 2.9 out of 5 versus 4.2 for short ones. Cost per long chat is high because the full transcript is re-sent every turn.

The team adds pinned slots (PNR, passenger names, flight dates, fare difference quoted), keeps 8 recent turns, and compacts at 60% of budget. On a 100-conversation recall eval, questions about turn-5 facts asked at turn 50 go from 58% correct to 97%. Tokens per turn fall by 55%, and satisfaction for long chats rises to 3.9.

Follow-up questions to expect

  • "Why not just use a million-token context model?" — Cost and latency grow with every turn, and models still use cluttered long context less reliably. Curation helps even when everything would fit.
  • "How is this different from long-term memory?" — This is within one session. Long-term memory stores facts across sessions and needs rules for going stale.
  • "What if compaction drops something important?" — The original turns stay in storage and can be retrieved; and anything that must never be lost belongs in a pinned slot.