Course Content
RAG Systems
12 sections · 66 lessons
What caching strategies are used in RAG pipelines?
What you need to know
Each cache layer catches a different kind of repeat.
| Layer | Key | Value | Risk |
|---|---|---|---|
| Embedding cache | hash(model + text) | vector | Almost none |
| Retrieval cache | normalised query + filters + index version | chunk ids | Stale after index updates |
| Exact response cache | normalised question + scope + versions | final answer | Stale answers |
| Semantic response cache | question embedding, similarity above threshold | final answer | Wrong answer for a similar question |
| Prompt cache (provider) | exact prompt prefix | model's internal state | None for correctness |
The semantic cache danger, with real numbers
A semantic cache returns a stored answer when the new question's embedding is close enough to a cached one. I measured cosine similarity with a common small embedding model (all-MiniLM-L6-v2):
| Question A | Question B | Cosine | Same answer? |
|---|---|---|---|
| What is the annual fee for the Platinum card? | Is there an annual fee on the Platinum card? | 0.951 | Yes |
| What is the annual fee for the Platinum card? | How much is the yearly charge on the Platinum credit card? | 0.915 | Yes |
| Can I return shoes after 30 days? | Can I return shoes after 10 days? | 0.942 | No |
| What is the annual fee for the Platinum card? | What is the annual fee for the Gold card? | 0.861 | No |
Look at rows 2 and 3. A true paraphrase scored 0.915, and a question with a different answer scored 0.942. No threshold separates them. Embeddings capture topic well and small details like numbers and product names poorly.
So, for semantic caching:
- Use it for high-volume, general questions ("how do I reset my password?"), not for questions with numbers, names or personal data.
- Add an exact-match check on key entities (product, number, date) extracted from both questions.
- Key it by tenant and the user's permission scope, so one user's cached answer never reaches someone who cannot see the source.
- Never cache personalised answers ("what is my leave balance?").
Invalidation
Put version numbers in the key: index_version, prompt_version, model_version. When you re-index or change the prompt, old entries are simply never matched again. Add a TTL (time to live) as a safety net.
Prompt caching
Providers can cache the processed form of a prompt prefix that repeats exactly: long system instructions, few-shot examples, a fixed reference document. Put the fixed part first and the retrieved chunks and question last. Cached input is billed at a reduced rate and processed faster. It is the only layer here with no correctness risk.
A real-life example
An e-commerce support bot adds a semantic cache with a 0.90 threshold before a big sale. Hit rate reaches 35%, and costs fall. Two days later, complaints arrive: customers asking about returning items "after 10 days" are told the answer for "after 30 days", and questions about one phone model get the answer for another.
The team changes three things: raise the threshold to 0.96, require that product ids and numbers extracted from both questions match exactly, and turn off semantic caching for order-specific questions. Hit rate drops to 18%, but a weekly sample of 200 cached answers now shows no wrong hits. They keep the embedding and prompt caches, which were never the problem.
Follow-up questions to expect
- "How do you measure a cache's quality?" — Hit rate for savings, plus a sampled wrong-hit rate: have a judge or a human check whether the cached answer really answers the new question.
- "Where do you store it?" — Redis or a similar store for exact keys; a small vector index for semantic keys, both with TTLs.
- "Does caching break access control?" — It can, if keys ignore permissions. Include tenant and permission scope in every key.