Retrieval-Augmented Generation (RAG)

Course Content

Retrieval-Augmented Generation (RAG)

4 sections · 8 lessons

Integrating Caching and Memory Systems


A subscription company added a semantic cache to their support RAG bot. The logic was sensible: if a new question is close enough to one already answered, reuse the stored answer instead of paying for retrieval and generation again. They set the similarity threshold at 0.92 and watched costs drop 31% in a week.

Three weeks later a customer asked "Can I get a refund after 30 days?" and the bot said "Yes, refunds are available on request."

That answer had been generated for a different question — "Can I get a refund?" — whose correct answer was indeed yes, within the 14-day window. The cosine similarity between the two questions was 0.94. The embedding model treated "after 30 days" as a mild qualifier on a refund question. In fact it inverted the answer completely.

Forty-one customers received that answer before anyone noticed. The refunds were honoured, because the company had said yes in writing. The cache had saved roughly 190 dollars and cost about 6,000.

That is the shape of every caching bug in a RAG system: the savings are immediate and easy to measure, the errors are delayed and hard to see. Memory has the same shape. Both are worth building. Both need to be built with the failure modes in view from the start.

Choosing a semantic cache threshold0.8562 percentabout 1 in 6no0.9044 percentabout 1 in 20risky0.9231 percentabout 1 in 40with logging0.9512 percentunder 1 in 500yes0.983 percenteffectively noneyes, low valueThresholdHit rateWrong answer servedShip itMeasure both columns on your own query log — the numbers move with the embedding model.
A cache hit is not a saving if the stored answer belongs to a different question, so the threshold is a correctness setting that happens to also reduce cost.

What caching actually buys you

Price out a single uncached RAG request before deciding whether any of this is worth the complexity.

StageLatencyCost
Embed the query40 msnegligible
Vector search (2M chunks, HNSW)15 msinfrastructure only
Cross-encoder re-rank (50 → 5)120 msinfrastructure only
Generation (4,000 in, 300 out)1,400 ms0.0165 dollars
Total1,575 ms0.0165 dollars

Generation is 89% of the latency and essentially all of the cost. At 500,000 queries a month that is 8,250 dollars and a p50 response that users describe as "a bit slow".

Now the fact that makes caching viable: real query traffic is not uniform. It follows a Zipf-like distribution. In most support corpora the 100 most common distinct questions account for 20-40% of all traffic. People ask how to reset their password, what the refund policy is, and how to add a seat — over and over, in slightly different words.

Caching is only interesting because users repeat themselves. Measure your repetition rate before you build anything: if the top 100 questions cover 4% of traffic, no cache design will help you.

Strategy 1 — exact query caching

The simplest useful cache. Normalise the query text, hash it, look it up.

Python
import hashlib, json, redef cache_key(query, tenant_id, acl_hash, cfg):    norm = re.sub(r"\s+", " ", query.lower().strip())    norm = re.sub(r"[^\w\s]", "", norm)    payload = json.dumps({        "q": norm,        "tenant": tenant_id,        # never omit        "acl": acl_hash,            # never omit        "index": cfg["index_version"],        "prompt": cfg["prompt_version"],        "model": cfg["model_id"],    }, sort_keys=True)    return "rag:v1:" + hashlib.sha256(payload.encode()).hexdigest()

Normalisation is what turns a useless cache into a working one. "How do I reset my password?", "how do i reset my password" and "How do I reset my password" are three distinct strings and one question. Lowercasing, collapsing whitespace and stripping punctuation typically doubles the hit rate, from around 6% to 12-18%.

Everything else in that key exists because of a specific way systems break.

Key componentWhat happens if you leave it out
tenantCustomer A asks "what are our payment terms", customer B asks the same, and B receives A's contract terms. This is a data breach, not a bug.
aclA manager's answer is served to a contractor, including the salary bands the contractor cannot see.
index_versionYou re-index the corpus with better chunking and the cache keeps serving answers built from the old chunks for as long as the TTL allows.
prompt_versionYou fix a prompt bug, deploy, and 30% of traffic keeps showing the bug.
model_idYou upgrade the generator and cannot tell whether quality improved, because a third of responses came from the old one.

The tenant and ACL entries are the ones that end careers. Cache keyed on query text alone is correct for exactly one class of system: a single-tenant, public corpus with no per-user permissions. Anything else needs the full key.

Strategy 2 — semantic caching

Exact caching misses "how can I reset my password" against "how do I reset my password". Semantic caching embeds the query and searches a small vector index of previously answered queries:

Python
def semantic_lookup(query, embedder, cache_index, threshold=0.95):    qv = embedder.embed_query(query)    hits = cache_index.search(qv, k=1)    if hits and hits[0].score >= threshold:        return hits[0].payload["answer"], hits[0].score    return None, hits[0].score if hits else 0.0

Everything now hinges on threshold, and the intuition most people bring to it is wrong.

Choosing the threshold with arithmetic, not feel

Look at what real cosine similarities mean:

Cached queryNew queryCosineSame answer?
how do I reset my passwordhow can I reset my password0.98Yes
is the API rate limitedwhat is the API rate limit0.96Yes
is data encrypted at restis data not encrypted at rest0.96No — inverted
can I get a refundcan I get a refund after 30 days0.94No — inverted
what is the refund windowwhat is the refund window for enterprise0.93No — different plan

Notice that the dangerous pairs sit at 0.93-0.96, right where a "sensible" threshold lands. Embeddings are near-blind to negation, and they treat numbers, dates and qualifying clauses as weak signals — exactly the tokens that most often flip an answer.

So price the decision. Take 1,000 queries, an uncached cost of 0.0165 dollars each, and assume a wrong answer costs 0.50 dollars in support handling and goodwill:

ThresholdHit rateFalse-hit rateSavedHarmNet
0.8541%12%6.7724.60−17.83
0.9028%4.5%4.626.30−1.68
0.9517%0.9%2.810.77+2.04
0.989%0.1%1.490.05+1.44

Working the 0.85 row: 410 hits saved at 0.0165 is 6.765 dollars; 12% of 410 is 49.2 wrong answers at 0.50 is 24.60 dollars. The loose threshold with the impressive hit rate loses money. The optimum here is 0.95 — barely above the exact-match rate, and worth 2.04 dollars per thousand queries.

The cache hit rate is a vanity metric. The number that matters is net value per thousand queries, and it turns negative long before the hit rate stops looking good.

Two adjustments make semantic caching much safer:

  • Never cache across a negation or numeric difference. Before accepting a hit, compare the numbers, dates and negation words in the two queries. If they differ at all, treat it as a miss. This is a ten-line check and it removes most of the residual false-hit rate.
  • Raise the threshold for high-stakes intents. Billing, cancellation, security and compliance questions get 0.99 or no cache at all. "What are your office hours" can sit at 0.92.

Strategy 3 — caching embeddings separately

Embeddings are deterministic: the same model and the same text always produce the same vector. That makes them the safest thing in the system to cache, because a stale embedding is impossible — only a changed model or changed text can invalidate one, and both are in the key.

Python
def embed_cached(text, model_id, embedder, store):    key = "emb:" + hashlib.sha256(        (model_id + "\x00" + text).encode()).hexdigest()    hit = store.get(key)    if hit is not None:        return hit    vec = embedder.embed_query(text)    store.set(key, vec)          # no TTL needed; it can never go stale    return vec

The payoff is largest at ingestion. A 2,000,000-chunk corpus at 400 tokens a chunk is 800 million tokens to embed. At a sustained million tokens a minute that is roughly thirteen hours of wall-clock time per full re-index.

Now run a nightly crawl where 2% of documents changed. Without an embedding cache: thirteen hours, every night. With one: 98% of chunks hash to an existing entry, and the job finishes in about sixteen minutes. The cache turns a re-index from a scheduled event into a routine one.

Note the limit, because people get this wrong. Changing the chunk size changes the chunk text, which changes the hash, which misses the cache entirely. The embedding cache protects you against unchanged content, not against re-chunking.

Cache freshness: expiry and invalidation

A RAG cache stores answers derived from documents. When a document changes, every cached answer built from it is potentially wrong — and unlike a stale web page, a stale policy answer can be actively harmful.

StrategyHow it worksStaleness windowUse when
Fixed TTLEvery entry expires after N secondsUp to NDefault; corpus changes slowly and predictably
Tiered TTLTTL by volatility class of the sourceVaries by classMixed corpus — pricing changes weekly, reference docs yearly
Event invalidationDocument update purges answers citing itNear zeroYou control ingestion and can emit change events
Version stampingKey includes index/prompt/model versionZero on deployAlways — combine with one of the above

Tiered TTLs are the cheap 80% solution:

Content classTTLReasoning
Pricing, availability, quotas1 hourChanges without warning; wrong answers are commercially damaging
Policy, terms, procedures24 hoursChanges on a release cadence
Product reference, API docs7 daysVersioned and announced
Historical, archived, regulatory filings30 daysImmutable in practice

Event invalidation needs one design decision made up front, and it is the one teams skip: store the source document IDs alongside every cached answer. Without them you cannot tell which cached answers a document update affects, and your only option is flushing everything.

Python
cache.set(key, {    "answer": answer,    "source_ids": [d.metadata["doc_id"] for d in docs],  # the crucial bit    "created_at": now(),})def on_document_changed(doc_id):    # Reverse index: doc_id -> set of cache keys that cited it    for key in reverse_index.pop(doc_id, []):        cache.delete(key)

Conversation memory: why RAG makes it harder

A plain chatbot with memory only has to keep the model coherent. A RAG system has a second problem, and it is the harder one.

Text
User:  What is the refund window?Bot:   14 days from the invoice date.User:  Does that apply to annual plans?

Send "Does that apply to annual plans?" to the retriever and it will search for documents about annual plans. It has no idea what "that" is. The chunk about the 14-day refund window scores poorly, because the query never mentions refunds. Retrieval fails, and then the generator either says something vague or invents an answer.

The naive fix — concatenate the whole history into the retrieval query — makes it worse. A ten-turn conversation that wandered through billing, SSO and rate limits produces a query vector pointing at the centroid of three unrelated topics, matching nothing well.

Query condensation is the fix

Rewrite the follow-up into a standalone question before it reaches the retriever.

Python
CONDENSE = """Given the conversation and a follow-up question, rewrite thefollow-up as a standalone question that makes sense with no prior context.Resolve every pronoun and implicit reference. Change nothing else.Conversation:{history}Follow-up: {question}Standalone question:"""def condense(history, question, small_llm):    if not history:        return question    return small_llm.invoke(        CONDENSE.format(history=render(history[-3:]), question=question)    ).content.strip()

That turns "Does that apply to annual plans?" into "Does the 14-day refund window apply to annual plans?" — a query the retriever can actually serve. It costs one small-model call, roughly 200 tokens and 120 ms, and on multi-turn traffic it is routinely worth twenty points of retrieval recall. It is the single highest-value component in conversational RAG.

In conversational RAG, memory's first job is not to remind the model what was said. It is to make the follow-up question retrievable.

How much history to carry

Unbounded history is quadratic. At about 150 tokens per turn and 4,000 tokens of retrieved context, a 40-turn conversation costs ∑k=039(4000+150k)=160,000+117,000=277,000\sum_{k=0}^{39}(4000 + 150k) = 160{,}000 + 117{,}000 = 277{,}000 input tokens — 0.83 dollars for one conversation, with the history share growing every turn. By turn 100 a single request carries 15,000 tokens of chat history against 4,000 tokens of retrieved documents, so the model is attending mostly to itself rather than to your corpus.

Memory typeTokens at turn 40KeepsLoses
Full buffer~6,000Everything, verbatimBounded cost; crowds out retrieved context
Sliding window (last 6 turns)~900Recent detailAnything established early
Running summary~300The gistExact figures and names; costs an LLM call per turn
Summary + window~1,200Gist of old, verbatim recentLittle in practice — the default choice
Vector-retrieved history~600Relevant old turns on demandChronology; adds a second retrieval

Summary plus window is the practical default: summarise everything older than the last six turns, keep those six verbatim. Numbers, IDs and dates survive because they are usually recent; the older material survives as gist, which is what it is for.

A note on the LangChain memory classes

Plenty of tutorials still show ConversationBufferMemory, ConversationSummaryMemory and ConversationChain. Those classes are legacy: in LangChain 1.x they are no longer in the main langchain package and survive only in langchain-classic for old code. They held state in a Python object inside the chain, which meant memory vanished on restart, could not be shared across processes, and had no clean way to branch or inspect a conversation.

The current approach separates state from the chain. LangGraph persists conversation state through a checkpointer, keyed by a thread_id you supply:

Python
from langgraph.graph import StateGraph, MessagesState, STARTfrom langgraph.checkpoint.postgres import PostgresSaverfrom langchain_core.messages import trim_messagesdef answer(state: MessagesState):    history = state["messages"][:-1]    question = state["messages"][-1].content    standalone = condense(history, question, small_llm)    docs = retrieve_and_rerank(standalone, store, reranker)    kept = trim_messages(history, max_tokens=900,                         token_counter=llm, strategy="last")    return {"messages": [llm.invoke(build_prompt(kept, docs, question))]}builder = StateGraph(MessagesState)builder.add_node("answer", answer)builder.add_edge(START, "answer")with PostgresSaver.from_conn_string(DB_URL) as saver:    saver.setup()    graph = builder.compile(checkpointer=saver)    graph.invoke(        {"messages": [("user", "Does that apply to annual plans?")]},        config={"configurable": {"thread_id": "user-8123"}},    )

Three things follow from this design and none of them were available before. State survives a restart, because it lives in Postgres rather than in memory. Any worker can serve any conversation, because the thread ID is the only thing that needs routing. And you can read a conversation's full history back out for debugging, which is how you find out why a bot said something odd three days ago.

Tracking which documents the conversation has seen

RAG-specific memory goes beyond messages. Store the document IDs retrieved on each turn and two useful behaviours become possible.

Avoid re-injecting. If a chunk was in the context two turns ago and is still in the recent history, re-sending it verbatim wastes tokens the retriever could spend on something new.

Detect topic changes. Compare the retrieved document sets across turns with Jaccard overlap:

J(A,B)=∣A∩B∣∣A∪B∣J(A,B) = \frac{|A \cap B|}{|A \cup B|}

Turn 5 retrieves {d12, d40, d41, d77, d90}; turn 6 retrieves {d12, d41, d77, d90, d91}. The intersection has 4 elements, the union has 6, so J=4/6=0.67J = 4/6 = 0.67 — plainly the same topic, so keep the history window. Turn 7 retrieves {d201, d202, d203, d204, d205}: intersection 0, union 10, J=0.00J = 0.00. The user has switched subject entirely.

Below roughly J=0.15J = 0.15, drop the history window and condense from the new question alone. This kills the most annoying multi-turn bug there is — the bot that keeps dragging the previous topic into a new one because "refund" appeared eight turns ago.

The correctness risks these two features introduce

Both caching and memory work by reusing something computed earlier. Every risk below is a variation on that reuse being invalid.

SymptomCauseFix
User sees another customer's dataCache key omits tenant or ACL hashInclude both in the key; test it deliberately with two tenants asking identical questions
Answer contradicts a document updated yesterdayNo invalidation path from document to cached answerStore source IDs per entry; maintain a reverse index; purge on change
Answer inverted — "yes" where the truth is "no"Semantic cache hit across a negation or a numeric qualifierThreshold at 0.95+, plus an explicit number/date/negation comparison before accepting
One user's personal detail appears in another's answerAn answer containing user-specific data was cached globallyNever cache when the prompt included profile data or the answer contains a name, date or account figure
Bot repeats and builds on an earlier mistakeMemory poisoning — a wrong answer in history is treated as established factKeep retrieved documents authoritative over history in the prompt; let users correct the record; expire long threads
Same question, two different answers, one frozen foreverA temperature > 0 sample was cachedCache only deterministic settings, or accept that the cache freezes one sample and set temperature to 0 for cached paths
Follow-ups retrieve nothing usefulNo query condensation; pronouns reach the retriever unresolvedCondense every follow-up to a standalone question before retrieval

The personalisation risk deserves the extra sentence. An answer like "Your Pro plan renews on 3 March and covers 12 seats" is perfectly correct and completely uncacheable. The safe rule is to cache only answers derived purely from corpus documents, with no user profile in the prompt — and to run a cheap regex for dates, currency amounts and proper nouns as a second gate before writing to the cache.

What this means when you build one

Do these in order, and stop as soon as the numbers stop justifying the next step.

Cache embeddings first. It cannot go stale, it cannot leak, it needs no threshold, and it turns re-indexing from a thirteen-hour job into a sixteen-minute one. There is no argument against it.

Then add an exact query cache with the full key — normalised query, tenant, ACL hash, index version, prompt version, model ID — and tiered TTLs. Expect 12-18% hit rate for a few hours of work, and no possibility of a false hit, because exact means exact.

Only then consider semantic caching, and only after you have measured your own false-hit rate on real traffic. Sample 200 cache hits, have a human check whether the cached answer was actually right for the new question, and put that number into the arithmetic above. If net value per thousand queries is not comfortably positive at a threshold of 0.95, do not ship it. Exclude billing, cancellation and security intents from the semantic cache entirely, whatever the numbers say — the average cost of a wrong answer is not the cost of the wrong answers that matter.

For memory, build query condensation before you build anything else. It is one small-model call and it fixes the failure that actually breaks conversational RAG: follow-up questions that the retriever cannot serve. Windowing, summarisation and persistence are all optimisations on top of a system that already retrieves correctly on turn two. Persist state through a checkpointer keyed by thread ID so it survives restarts and any worker can pick up any conversation, keep six turns verbatim with a running summary behind them, and track retrieved document IDs so you can tell when the subject has changed and clear the window.

Then instrument the thing you cannot see. Log every cache hit with its similarity score and the query it matched, and sample a hundred of them a week for manual review. The refund bug in the opening was visible in the logs for nineteen days before a human read them.