Course Content
Agentic AI Patterns
9 sections · 50 lessons
How does memory optimization improve model performance in Agentic AI systems?
What you need to know
Why context grows so fast in agents
An agent re-sends the whole conversation on every step. Take a 12-step run where the system prompt and tool definitions are 6,000 tokens and each step adds about 1,800 tokens:
Step 1 input: 6,000 tokensStep 12 input: 6,000 + 11 x 1,800 = 25,800 tokensTotal input over 12 steps: about 190,800 tokensOutput: 12 x 300 = 3,600 tokensCost at an assumed $3 / $15 per million in / out: about $0.57 + $0.05 = $0.63At 10,000 runs a day, that is about $6,300 a day.
The techniques
- Structured state. Keep a typed object (goal, plan, findings, open questions) and render it each step, instead of replaying every message.
- Compaction. When the transcript passes a threshold, summarise the older part and keep recent turns word for word.
- Tool-result discipline. Store the full result under an ID; put a short version in context. Raw tool output is usually the largest item.
- Retrieval over stuffing. Fetch the 3 to 5 relevant memories, not all of them.
- Prompt caching. Most of each step's input is identical to the previous step. With caching, repeated prefix tokens are billed at a steep discount. In the run above, assuming cached reads at a tenth of the price and a small surcharge on newly cached tokens, the cost drops from about $0.63 to about $0.20.
- Deduplication and decay in long-term stores.
Why it improves quality
Models do not use long contexts evenly. The "lost in the middle" finding showed that facts placed in the middle of a long prompt are used less reliably than those at the start or end. Noise also misleads: an agent that sees three stale error messages may retry a tool that already recovered. Cleaner context often raises the success rate outright.
A real-life example
An insurance-claims agent handled complex claims with 20 to 30 steps. Each step appended the full get_claim_documents result, including OCR text of every page. By step 15 the context was 70,000 tokens, and the agent started asking for documents it had already received.
Changes the team made:
- Documents are stored by ID; context gets a 3-line summary per document plus the extracted fields.
- A state object tracks
documents_received,documents_missing,checks_done. - The system prompt and tool list were moved to the front and kept byte-identical, so caching works.
Average context at step 15 fell to about 14,000 tokens. Repeated document requests stopped, and success on the golden set rose from 71% to 79%. Cost per successful claim fell by roughly 60%.
Follow-up questions to expect
- "What is the risk of summarising?" — Losing a detail that matters later. Pin critical facts, like the claim amount, in the state object so summaries can never drop them.
- "How does prompt caching work?" — The provider reuses the processed prefix of a prompt that matches a recent request exactly. Any change early in the prompt breaks the cache for everything after it.
- "What metric shows the optimisation worked?" — Cost per successful task and success rate together. Cheaper but less successful is not an improvement.