Course Content
Retrieval-Augmented Generation (RAG)
4 sections · 8 lessons
Integrating Caching and Memory Systems
A subscription company added a semantic cache to their support RAG bot. The logic was sensible: if a new question is close enough to one already answered, reuse the stored answer instead of paying for retrieval and generation again. They set the similarity threshold at 0.92 and watched costs drop 31% in a week.
Three weeks later a customer asked "Can I get a refund after 30 days?" and the bot said "Yes, refunds are available on request."
That answer had been generated for a different question — "Can I get a refund?" — whose correct answer was indeed yes, within the 14-day window. The cosine similarity between the two questions was 0.94. The embedding model treated "after 30 days" as a mild qualifier on a refund question. In fact it inverted the answer completely.
Forty-one customers received that answer before anyone noticed. The refunds were honoured, because the company had said yes in writing. The cache had saved roughly 190 dollars and cost about 6,000.
That is the shape of every caching bug in a RAG system: the savings are immediate and easy to measure, the errors are delayed and hard to see. Memory has the same shape. Both are worth building. Both need to be built with the failure modes in view from the start.
What caching actually buys you
Price out a single uncached RAG request before deciding whether any of this is worth the complexity.
| Stage | Latency | Cost |
|---|---|---|
| Embed the query | 40 ms | negligible |
| Vector search (2M chunks, HNSW) | 15 ms | infrastructure only |
| Cross-encoder re-rank (50 → 5) | 120 ms | infrastructure only |
| Generation (4,000 in, 300 out) | 1,400 ms | 0.0165 dollars |
| Total | 1,575 ms | 0.0165 dollars |
Generation is 89% of the latency and essentially all of the cost. At 500,000 queries a month that is 8,250 dollars and a p50 response that users describe as "a bit slow".
Now the fact that makes caching viable: real query traffic is not uniform. It follows a Zipf-like distribution. In most support corpora the 100 most common distinct questions account for 20-40% of all traffic. People ask how to reset their password, what the refund policy is, and how to add a seat — over and over, in slightly different words.
Caching is only interesting because users repeat themselves. Measure your repetition rate before you build anything: if the top 100 questions cover 4% of traffic, no cache design will help you.
Strategy 1 — exact query caching
The simplest useful cache. Normalise the query text, hash it, look it up.
1import hashlib, json, re23def cache_key(query, tenant_id, acl_hash, cfg):4 norm = re.sub(r"\s+", " ", query.lower().strip())5 norm = re.sub(r"[^\w\s]", "", norm)6 payload = json.dumps({7 "q": norm,8 "tenant": tenant_id, # never omit9 "acl": acl_hash, # never omit10 "index": cfg["index_version"],11 "prompt": cfg["prompt_version"],12 "model": cfg["model_id"],13 }, sort_keys=True)14 return "rag:v1:" + hashlib.sha256(payload.encode()).hexdigest()Normalisation is what turns a useless cache into a working one. "How do I reset my password?", "how do i reset my password" and "How do I reset my password" are three distinct strings and one question. Lowercasing, collapsing whitespace and stripping punctuation typically doubles the hit rate, from around 6% to 12-18%.
Everything else in that key exists because of a specific way systems break.
| Key component | What happens if you leave it out |
|---|---|
tenant | Customer A asks "what are our payment terms", customer B asks the same, and B receives A's contract terms. This is a data breach, not a bug. |
acl | A manager's answer is served to a contractor, including the salary bands the contractor cannot see. |
index_version | You re-index the corpus with better chunking and the cache keeps serving answers built from the old chunks for as long as the TTL allows. |
prompt_version | You fix a prompt bug, deploy, and 30% of traffic keeps showing the bug. |
model_id | You upgrade the generator and cannot tell whether quality improved, because a third of responses came from the old one. |
The tenant and ACL entries are the ones that end careers. Cache keyed on query text alone is correct for exactly one class of system: a single-tenant, public corpus with no per-user permissions. Anything else needs the full key.
Strategy 2 — semantic caching
Exact caching misses "how can I reset my password" against "how do I reset my password". Semantic caching embeds the query and searches a small vector index of previously answered queries:
1def semantic_lookup(query, embedder, cache_index, threshold=0.95):2 qv = embedder.embed_query(query)3 hits = cache_index.search(qv, k=1)4 if hits and hits[0].score >= threshold:5 return hits[0].payload["answer"], hits[0].score6 return None, hits[0].score if hits else 0.0Everything now hinges on threshold, and the intuition most people bring to it is wrong.
Choosing the threshold with arithmetic, not feel
Look at what real cosine similarities mean:
| Cached query | New query | Cosine | Same answer? |
|---|---|---|---|
| how do I reset my password | how can I reset my password | 0.98 | Yes |
| is the API rate limited | what is the API rate limit | 0.96 | Yes |
| is data encrypted at rest | is data not encrypted at rest | 0.96 | No — inverted |
| can I get a refund | can I get a refund after 30 days | 0.94 | No — inverted |
| what is the refund window | what is the refund window for enterprise | 0.93 | No — different plan |
Notice that the dangerous pairs sit at 0.93-0.96, right where a "sensible" threshold lands. Embeddings are near-blind to negation, and they treat numbers, dates and qualifying clauses as weak signals — exactly the tokens that most often flip an answer.
So price the decision. Take 1,000 queries, an uncached cost of 0.0165 dollars each, and assume a wrong answer costs 0.50 dollars in support handling and goodwill:
| Threshold | Hit rate | False-hit rate | Saved | Harm | Net |
|---|---|---|---|---|---|
| 0.85 | 41% | 12% | 6.77 | 24.60 | −17.83 |
| 0.90 | 28% | 4.5% | 4.62 | 6.30 | −1.68 |
| 0.95 | 17% | 0.9% | 2.81 | 0.77 | +2.04 |
| 0.98 | 9% | 0.1% | 1.49 | 0.05 | +1.44 |
Working the 0.85 row: 410 hits saved at 0.0165 is 6.765 dollars; 12% of 410 is 49.2 wrong answers at 0.50 is 24.60 dollars. The loose threshold with the impressive hit rate loses money. The optimum here is 0.95 — barely above the exact-match rate, and worth 2.04 dollars per thousand queries.
The cache hit rate is a vanity metric. The number that matters is net value per thousand queries, and it turns negative long before the hit rate stops looking good.
Two adjustments make semantic caching much safer:
- Never cache across a negation or numeric difference. Before accepting a hit, compare the numbers, dates and negation words in the two queries. If they differ at all, treat it as a miss. This is a ten-line check and it removes most of the residual false-hit rate.
- Raise the threshold for high-stakes intents. Billing, cancellation, security and compliance questions get 0.99 or no cache at all. "What are your office hours" can sit at 0.92.
Strategy 3 — caching embeddings separately
Embeddings are deterministic: the same model and the same text always produce the same vector. That makes them the safest thing in the system to cache, because a stale embedding is impossible — only a changed model or changed text can invalidate one, and both are in the key.
1def embed_cached(text, model_id, embedder, store):2 key = "emb:" + hashlib.sha256(3 (model_id + "\x00" + text).encode()).hexdigest()4 hit = store.get(key)5 if hit is not None:6 return hit7 vec = embedder.embed_query(text)8 store.set(key, vec) # no TTL needed; it can never go stale9 return vecThe payoff is largest at ingestion. A 2,000,000-chunk corpus at 400 tokens a chunk is 800 million tokens to embed. At a sustained million tokens a minute that is roughly thirteen hours of wall-clock time per full re-index.
Now run a nightly crawl where 2% of documents changed. Without an embedding cache: thirteen hours, every night. With one: 98% of chunks hash to an existing entry, and the job finishes in about sixteen minutes. The cache turns a re-index from a scheduled event into a routine one.
Note the limit, because people get this wrong. Changing the chunk size changes the chunk text, which changes the hash, which misses the cache entirely. The embedding cache protects you against unchanged content, not against re-chunking.
Cache freshness: expiry and invalidation
A RAG cache stores answers derived from documents. When a document changes, every cached answer built from it is potentially wrong — and unlike a stale web page, a stale policy answer can be actively harmful.
| Strategy | How it works | Staleness window | Use when |
|---|---|---|---|
| Fixed TTL | Every entry expires after N seconds | Up to N | Default; corpus changes slowly and predictably |
| Tiered TTL | TTL by volatility class of the source | Varies by class | Mixed corpus — pricing changes weekly, reference docs yearly |
| Event invalidation | Document update purges answers citing it | Near zero | You control ingestion and can emit change events |
| Version stamping | Key includes index/prompt/model version | Zero on deploy | Always — combine with one of the above |
Tiered TTLs are the cheap 80% solution:
| Content class | TTL | Reasoning |
|---|---|---|
| Pricing, availability, quotas | 1 hour | Changes without warning; wrong answers are commercially damaging |
| Policy, terms, procedures | 24 hours | Changes on a release cadence |
| Product reference, API docs | 7 days | Versioned and announced |
| Historical, archived, regulatory filings | 30 days | Immutable in practice |
Event invalidation needs one design decision made up front, and it is the one teams skip: store the source document IDs alongside every cached answer. Without them you cannot tell which cached answers a document update affects, and your only option is flushing everything.
1cache.set(key, {2 "answer": answer,3 "source_ids": [d.metadata["doc_id"] for d in docs], # the crucial bit4 "created_at": now(),5})67def on_document_changed(doc_id):8 # Reverse index: doc_id -> set of cache keys that cited it9 for key in reverse_index.pop(doc_id, []):10 cache.delete(key)Conversation memory: why RAG makes it harder
A plain chatbot with memory only has to keep the model coherent. A RAG system has a second problem, and it is the harder one.
User: What is the refund window?Bot: 14 days from the invoice date.User: Does that apply to annual plans?Send "Does that apply to annual plans?" to the retriever and it will search for documents about annual plans. It has no idea what "that" is. The chunk about the 14-day refund window scores poorly, because the query never mentions refunds. Retrieval fails, and then the generator either says something vague or invents an answer.
The naive fix — concatenate the whole history into the retrieval query — makes it worse. A ten-turn conversation that wandered through billing, SSO and rate limits produces a query vector pointing at the centroid of three unrelated topics, matching nothing well.
Query condensation is the fix
Rewrite the follow-up into a standalone question before it reaches the retriever.
1CONDENSE = """Given the conversation and a follow-up question, rewrite the2follow-up as a standalone question that makes sense with no prior context.3Resolve every pronoun and implicit reference. Change nothing else.45Conversation:6{history}78Follow-up: {question}9Standalone question:"""1011def condense(history, question, small_llm):12 if not history:13 return question14 return small_llm.invoke(15 CONDENSE.format(history=render(history[-3:]), question=question)16 ).content.strip()That turns "Does that apply to annual plans?" into "Does the 14-day refund window apply to annual plans?" — a query the retriever can actually serve. It costs one small-model call, roughly 200 tokens and 120 ms, and on multi-turn traffic it is routinely worth twenty points of retrieval recall. It is the single highest-value component in conversational RAG.
In conversational RAG, memory's first job is not to remind the model what was said. It is to make the follow-up question retrievable.
How much history to carry
Unbounded history is quadratic. At about 150 tokens per turn and 4,000 tokens of retrieved context, a 40-turn conversation costs ∑k=039(4000+150k)=160,000+117,000=277,000 input tokens — 0.83 dollars for one conversation, with the history share growing every turn. By turn 100 a single request carries 15,000 tokens of chat history against 4,000 tokens of retrieved documents, so the model is attending mostly to itself rather than to your corpus.
| Memory type | Tokens at turn 40 | Keeps | Loses |
|---|---|---|---|
| Full buffer | ~6,000 | Everything, verbatim | Bounded cost; crowds out retrieved context |
| Sliding window (last 6 turns) | ~900 | Recent detail | Anything established early |
| Running summary | ~300 | The gist | Exact figures and names; costs an LLM call per turn |
| Summary + window | ~1,200 | Gist of old, verbatim recent | Little in practice — the default choice |
| Vector-retrieved history | ~600 | Relevant old turns on demand | Chronology; adds a second retrieval |
Summary plus window is the practical default: summarise everything older than the last six turns, keep those six verbatim. Numbers, IDs and dates survive because they are usually recent; the older material survives as gist, which is what it is for.
A note on the LangChain memory classes
Plenty of tutorials still show ConversationBufferMemory, ConversationSummaryMemory and ConversationChain. Those classes are legacy: in LangChain 1.x they are no longer in the main langchain package and survive only in langchain-classic for old code. They held state in a Python object inside the chain, which meant memory vanished on restart, could not be shared across processes, and had no clean way to branch or inspect a conversation.
The current approach separates state from the chain. LangGraph persists conversation state through a checkpointer, keyed by a thread_id you supply:
1from langgraph.graph import StateGraph, MessagesState, START2from langgraph.checkpoint.postgres import PostgresSaver3from langchain_core.messages import trim_messages45def answer(state: MessagesState):6 history = state["messages"][:-1]7 question = state["messages"][-1].content8 standalone = condense(history, question, small_llm)9 docs = retrieve_and_rerank(standalone, store, reranker)10 kept = trim_messages(history, max_tokens=900,11 token_counter=llm, strategy="last")12 return {"messages": [llm.invoke(build_prompt(kept, docs, question))]}1314builder = StateGraph(MessagesState)15builder.add_node("answer", answer)16builder.add_edge(START, "answer")1718with PostgresSaver.from_conn_string(DB_URL) as saver:19 saver.setup()20 graph = builder.compile(checkpointer=saver)21 graph.invoke(22 {"messages": [("user", "Does that apply to annual plans?")]},23 config={"configurable": {"thread_id": "user-8123"}},24 )Three things follow from this design and none of them were available before. State survives a restart, because it lives in Postgres rather than in memory. Any worker can serve any conversation, because the thread ID is the only thing that needs routing. And you can read a conversation's full history back out for debugging, which is how you find out why a bot said something odd three days ago.
Tracking which documents the conversation has seen
RAG-specific memory goes beyond messages. Store the document IDs retrieved on each turn and two useful behaviours become possible.
Avoid re-injecting. If a chunk was in the context two turns ago and is still in the recent history, re-sending it verbatim wastes tokens the retriever could spend on something new.
Detect topic changes. Compare the retrieved document sets across turns with Jaccard overlap:
Turn 5 retrieves {d12, d40, d41, d77, d90}; turn 6 retrieves {d12, d41, d77, d90, d91}. The intersection has 4 elements, the union has 6, so J=4/6=0.67 — plainly the same topic, so keep the history window. Turn 7 retrieves {d201, d202, d203, d204, d205}: intersection 0, union 10, J=0.00. The user has switched subject entirely.
Below roughly J=0.15, drop the history window and condense from the new question alone. This kills the most annoying multi-turn bug there is — the bot that keeps dragging the previous topic into a new one because "refund" appeared eight turns ago.
The correctness risks these two features introduce
Both caching and memory work by reusing something computed earlier. Every risk below is a variation on that reuse being invalid.
| Symptom | Cause | Fix |
|---|---|---|
| User sees another customer's data | Cache key omits tenant or ACL hash | Include both in the key; test it deliberately with two tenants asking identical questions |
| Answer contradicts a document updated yesterday | No invalidation path from document to cached answer | Store source IDs per entry; maintain a reverse index; purge on change |
| Answer inverted — "yes" where the truth is "no" | Semantic cache hit across a negation or a numeric qualifier | Threshold at 0.95+, plus an explicit number/date/negation comparison before accepting |
| One user's personal detail appears in another's answer | An answer containing user-specific data was cached globally | Never cache when the prompt included profile data or the answer contains a name, date or account figure |
| Bot repeats and builds on an earlier mistake | Memory poisoning — a wrong answer in history is treated as established fact | Keep retrieved documents authoritative over history in the prompt; let users correct the record; expire long threads |
| Same question, two different answers, one frozen forever | A temperature > 0 sample was cached | Cache only deterministic settings, or accept that the cache freezes one sample and set temperature to 0 for cached paths |
| Follow-ups retrieve nothing useful | No query condensation; pronouns reach the retriever unresolved | Condense every follow-up to a standalone question before retrieval |
The personalisation risk deserves the extra sentence. An answer like "Your Pro plan renews on 3 March and covers 12 seats" is perfectly correct and completely uncacheable. The safe rule is to cache only answers derived purely from corpus documents, with no user profile in the prompt — and to run a cheap regex for dates, currency amounts and proper nouns as a second gate before writing to the cache.
What this means when you build one
Do these in order, and stop as soon as the numbers stop justifying the next step.
Cache embeddings first. It cannot go stale, it cannot leak, it needs no threshold, and it turns re-indexing from a thirteen-hour job into a sixteen-minute one. There is no argument against it.
Then add an exact query cache with the full key — normalised query, tenant, ACL hash, index version, prompt version, model ID — and tiered TTLs. Expect 12-18% hit rate for a few hours of work, and no possibility of a false hit, because exact means exact.
Only then consider semantic caching, and only after you have measured your own false-hit rate on real traffic. Sample 200 cache hits, have a human check whether the cached answer was actually right for the new question, and put that number into the arithmetic above. If net value per thousand queries is not comfortably positive at a threshold of 0.95, do not ship it. Exclude billing, cancellation and security intents from the semantic cache entirely, whatever the numbers say — the average cost of a wrong answer is not the cost of the wrong answers that matter.
For memory, build query condensation before you build anything else. It is one small-model call and it fixes the failure that actually breaks conversational RAG: follow-up questions that the retriever cannot serve. Windowing, summarisation and persistence are all optimisations on top of a system that already retrieves correctly on turn two. Persist state through a checkpointer keyed by thread ID so it survives restarts and any worker can pick up any conversation, keep six turns verbatim with a running summary behind them, and track retrieved document IDs so you can tell when the subject has changed and clear the window.
Then instrument the thing you cannot see. Log every cache hit with its similarity score and the query it matched, and sample a hundred of them a week for manual review. The refund bug in the opening was visible in the logs for nineteen days before a human read them.