Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
By message 30, your chatbot context hits 40K tokens and each reply costs $0.40. Users churn due to slow responses. How do you compress conversation history without losing critical context?
What you need to know
Every turn re-sends everything before it. If each turn adds about 1,300 tokens, turn 30 sends about 40,000, and the conversation as a whole has sent the sum of all turns, roughly 600,000 tokens. Cost grows with the square of the length, and so does the time spent reading the prompt before the first word appears.
The three tiers
- Recent turns, verbatim — the last 6 to 10 messages. Conversational flow lives here; compressing it is what makes bots feel forgetful.
- Rolling summary — when the buffer passes a threshold, a small model summarises the oldest block and folds it into the existing summary. One cheap call every few turns, not a full re-summary.
- Structured facts — name, order id, account type, stated constraints ("vegetarian", "budget under ₹5,000"), decisions made. Stored as key-value state, always injected exactly.
1from langchain_core.messages import trim_messages23recent = trim_messages(history, max_tokens=6000, strategy="last",4 token_counter=llm, include_system=True, start_on="human")5prompt = [system_msg, facts_block(state.facts), summary_block(state.summary), *recent]trim_messages keeps the newest messages within a token budget; start_on="human" avoids starting mid tool-call. The facts and summary blocks are built from state the application keeps, not from the model's memory.
Why facts get their own tier
A summary will happily turn "order #48212, delivered to the Koramangala address" into "the user had a delivery issue". Structured extraction keeps the exact values the rest of the conversation depends on.
Options compared
| Strategy | Token cost | Keeps exact details? | Good for |
|---|---|---|---|
| Full history | Grows every turn | Yes | Short chats only |
| Last-N window | Fixed | Only recent ones | Casual chat |
| Rolling summary | Small and bounded | No, paraphrases | Long, topical conversations |
| Structured facts | Tiny | Yes, by design | Ids, preferences, decisions |
| Retrieval over old turns | Small | Yes, when retrieved | Topic jumps back to something from 50 turns ago |
Combine the first three tiers; add retrieval over archived turns if users often jump back to earlier topics.
The effect
A 40K-token turn falls to roughly 6 to 8K. At the same price per token, $0.40 becomes around $0.07, and time-to-first-token improves because prefill shrinks. Prompt caching on the stable system prefix cuts cost further.
A real-life example
Scenario (illustrative numbers). A travel company's trip-planning chatbot has long conversations: 40 to 60 turns while users compare hotels and trains. At turn 30, replies take 6 seconds to start and cost about $0.40 each. Session abandonment after turn 20 is 38%.
The team keeps the last 8 messages, a rolling summary, and a facts block with dates, budget, cities and the chosen hotel. Median prompt size at turn 30 falls to 7,200 tokens; time-to-first-token drops to 1.4 seconds. On a test set of 40 long transcripts with questions about early decisions ("which hotel did I pick for Jaipur?"), accuracy is 95%, compared with 97% for full history. Abandonment after turn 20 falls to 24%.
Follow-up questions to expect
- "What if the summariser drops something important?" — That is why facts and decisions are extracted separately. Test it with transcripts whose answers live in the compressed region, on every change.
- "Why not just use a model with a longer context?" — It still pays for every token every turn and still slows the first token; long context raises the ceiling but does not fix the cost curve.
- "When do you run the summariser?" — Asynchronously, after a reply is sent, so it adds no latency to the user's turn.