Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

By message 30, your chatbot context hits 40K tokens and each reply costs $0.40. Users churn due to slow responses. How do you compress conversation history without losing critical context?


What the prompt holds at turn 30Systemprompt, cachedFacts: dates,budget, hotelRolling summaryof turns 1-22Last 8messages verbatimtopbottomAbout 7,200 tokens instead of 40,000.
Identifiers live in their own block because a summary will quietly turn an order number into "an order".

What you need to know

Every turn re-sends everything before it. If each turn adds about 1,300 tokens, turn 30 sends about 40,000, and the conversation as a whole has sent the sum of all turns, roughly 600,000 tokens. Cost grows with the square of the length, and so does the time spent reading the prompt before the first word appears.

The three tiers

  1. Recent turns, verbatim — the last 6 to 10 messages. Conversational flow lives here; compressing it is what makes bots feel forgetful.
  2. Rolling summary — when the buffer passes a threshold, a small model summarises the oldest block and folds it into the existing summary. One cheap call every few turns, not a full re-summary.
  3. Structured facts — name, order id, account type, stated constraints ("vegetarian", "budget under ₹5,000"), decisions made. Stored as key-value state, always injected exactly.
Python
from langchain_core.messages import trim_messagesrecent = trim_messages(history, max_tokens=6000, strategy="last",                       token_counter=llm, include_system=True, start_on="human")prompt = [system_msg, facts_block(state.facts), summary_block(state.summary), *recent]

trim_messages keeps the newest messages within a token budget; start_on="human" avoids starting mid tool-call. The facts and summary blocks are built from state the application keeps, not from the model's memory.

Why facts get their own tier

A summary will happily turn "order #48212, delivered to the Koramangala address" into "the user had a delivery issue". Structured extraction keeps the exact values the rest of the conversation depends on.

Options compared

StrategyToken costKeeps exact details?Good for
Full historyGrows every turnYesShort chats only
Last-N windowFixedOnly recent onesCasual chat
Rolling summarySmall and boundedNo, paraphrasesLong, topical conversations
Structured factsTinyYes, by designIds, preferences, decisions
Retrieval over old turnsSmallYes, when retrievedTopic jumps back to something from 50 turns ago

Combine the first three tiers; add retrieval over archived turns if users often jump back to earlier topics.

The effect

A 40K-token turn falls to roughly 6 to 8K. At the same price per token, $0.40 becomes around $0.07, and time-to-first-token improves because prefill shrinks. Prompt caching on the stable system prefix cuts cost further.

A real-life example

Scenario (illustrative numbers). A travel company's trip-planning chatbot has long conversations: 40 to 60 turns while users compare hotels and trains. At turn 30, replies take 6 seconds to start and cost about $0.40 each. Session abandonment after turn 20 is 38%.

The team keeps the last 8 messages, a rolling summary, and a facts block with dates, budget, cities and the chosen hotel. Median prompt size at turn 30 falls to 7,200 tokens; time-to-first-token drops to 1.4 seconds. On a test set of 40 long transcripts with questions about early decisions ("which hotel did I pick for Jaipur?"), accuracy is 95%, compared with 97% for full history. Abandonment after turn 20 falls to 24%.

Follow-up questions to expect

  • "What if the summariser drops something important?" — That is why facts and decisions are extracted separately. Test it with transcripts whose answers live in the compressed region, on every change.
  • "Why not just use a model with a longer context?" — It still pays for every token every turn and still slows the first token; long context raises the ceiling but does not fix the cost curve.
  • "When do you run the summariser?" — Asynchronously, after a reply is sent, so it adds no latency to the user's turn.