RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

What caching strategies are used in RAG pipelines?


Measured cosine similarity against a cached question0.9510.9420.9150.861012330 days vs 10 days— different answertrue paraphraseGold vs Platinum —different answerall-MiniLM-L6-v2. A wrong match scores above a true paraphrase.
No single threshold separates these, so a semantic cache also needs exact checks on numbers and product names.

What you need to know

Each cache layer catches a different kind of repeat.

LayerKeyValueRisk
Embedding cachehash(model + text)vectorAlmost none
Retrieval cachenormalised query + filters + index versionchunk idsStale after index updates
Exact response cachenormalised question + scope + versionsfinal answerStale answers
Semantic response cachequestion embedding, similarity above thresholdfinal answerWrong answer for a similar question
Prompt cache (provider)exact prompt prefixmodel's internal stateNone for correctness

The semantic cache danger, with real numbers

A semantic cache returns a stored answer when the new question's embedding is close enough to a cached one. I measured cosine similarity with a common small embedding model (all-MiniLM-L6-v2):

Question AQuestion BCosineSame answer?
What is the annual fee for the Platinum card?Is there an annual fee on the Platinum card?0.951Yes
What is the annual fee for the Platinum card?How much is the yearly charge on the Platinum credit card?0.915Yes
Can I return shoes after 30 days?Can I return shoes after 10 days?0.942No
What is the annual fee for the Platinum card?What is the annual fee for the Gold card?0.861No

Look at rows 2 and 3. A true paraphrase scored 0.915, and a question with a different answer scored 0.942. No threshold separates them. Embeddings capture topic well and small details like numbers and product names poorly.

So, for semantic caching:

  • Use it for high-volume, general questions ("how do I reset my password?"), not for questions with numbers, names or personal data.
  • Add an exact-match check on key entities (product, number, date) extracted from both questions.
  • Key it by tenant and the user's permission scope, so one user's cached answer never reaches someone who cannot see the source.
  • Never cache personalised answers ("what is my leave balance?").

Invalidation

Put version numbers in the key: index_version, prompt_version, model_version. When you re-index or change the prompt, old entries are simply never matched again. Add a TTL (time to live) as a safety net.

Prompt caching

Providers can cache the processed form of a prompt prefix that repeats exactly: long system instructions, few-shot examples, a fixed reference document. Put the fixed part first and the retrieved chunks and question last. Cached input is billed at a reduced rate and processed faster. It is the only layer here with no correctness risk.

A real-life example

An e-commerce support bot adds a semantic cache with a 0.90 threshold before a big sale. Hit rate reaches 35%, and costs fall. Two days later, complaints arrive: customers asking about returning items "after 10 days" are told the answer for "after 30 days", and questions about one phone model get the answer for another.

The team changes three things: raise the threshold to 0.96, require that product ids and numbers extracted from both questions match exactly, and turn off semantic caching for order-specific questions. Hit rate drops to 18%, but a weekly sample of 200 cached answers now shows no wrong hits. They keep the embedding and prompt caches, which were never the problem.

Follow-up questions to expect

  • "How do you measure a cache's quality?" — Hit rate for savings, plus a sampled wrong-hit rate: have a judge or a human check whether the cached answer really answers the new question.
  • "Where do you store it?" — Redis or a similar store for exact keys; a small vector index for semantic keys, both with TTLs.
  • "Does caching break access control?" — It can, if keys ignore permissions. Include tenant and permission scope in every key.