Advanced RAG

Course Content

Advanced RAG

3 sections · 38 lessons

What is semantic caching, and how does it differ from simple caching?


A cache hit must pass two gatesEmbed thecondensed querySearch thisscope: tenant,ACL, languageSimilarity atleast 0.95?Numbers inquery matchcached question?Serve cachedanswer, log the hitA ₹399 question must never get the cached ₹299 answer.
The scope stops leaks between users; the threshold and number check stop near-identical questions getting each other's answers.

What you need to know

Exact-match cache

  • Key: the normalised query string
  • Hit only on identical text
  • Never returns a wrong answer for the text
  • Low hit rate in conversational products

Semantic cache

  • Key: the query's embedding, within a scope
  • Hit on paraphrases above a similarity threshold
  • Can return an answer to a different question
  • Much higher hit rate on FAQ-style traffic

How a lookup works

Python
import numpy as npclass SemanticCache:    def __init__(self, threshold=0.95):        self.threshold, self.entries = threshold, {}   # scope -> [(vec, answer)]    @staticmethod    def scope(tenant, acl_groups, lang, index_version, prompt_version):        return (tenant, tuple(sorted(acl_groups)), lang, index_version, prompt_version)    def get(self, scope, vec):        best, best_sim = None, -1.0        for v, answer in self.entries.get(scope, []):   # only search this scope            sim = float(v @ vec)            if sim > best_sim:                best, best_sim = answer, round(sim, 3)        return (best, best_sim) if best_sim >= self.threshold else (None, best_sim)    def put(self, scope, vec, answer):        self.entries.setdefault(scope, []).append((vec, answer))unit = lambda v: np.array(v) / np.linalg.norm(v)cache = SemanticCache()s = SemanticCache.scope("telco", ["prepaid"], "hi", "idx-42", "p-7")cache.put(s, unit([1.0, 0.2, 0.0]), "Dial *121# to check balance.")print(cache.get(s, unit([1.0, 0.25, 0.0])))      # ('Dial *121# ...', 0.999) hitprint(cache.get(s, unit([1.0, 0.6, 0.0])))       # (None, 0.942) similar, not enoughother = SemanticCache.scope("telco", ["postpaid"], "hi", "idx-42", "p-7")print(cache.get(other, unit([1.0, 0.2, 0.0])))   # (None, -1.0) other scope: no leak

Two ideas are in this code. The threshold decides hit or miss: 0.942 is "similar" but below 0.95, so it misses. The scope is part of the key: the same question from a postpaid user never sees the prepaid answer. In production the linear scan is replaced by a vector index, and tools such as Redis-based semantic caches or GPTCache provide this pattern.

The traps

  • The threshold is a precision dial. "How do I cancel my subscription?" and "How do I cancel my free trial?" are very close vectors with different answers. Tune on real traffic; a wrong cached answer is worse than a slow correct one.
  • Negation and small swaps are nearly invisible to embeddings: "Is roaming included?" versus "Is roaming not included?", "₹299 plan" versus "₹399 plan". Some teams add an exact check on numbers and entity names before accepting a hit.
  • Scope everything that changes the answer: tenant, user role or ACL groups, language, index version, prompt version, model version. Missing the tenant or ACL scope is a data leak, not a quality bug.
  • Freshness: set a TTL and invalidate entries when the documents they cited change.
  • Personal questions ("why is my bill high?") should never be cached.

Semantic cache versus prompt cache

A prompt cache (provider-side) reuses computation for a repeated prompt prefix; the model still generates a fresh answer. A semantic cache skips the model entirely and returns a stored answer. They solve different problems and work well together.

A real-life example

An Indian telecom's support bot gets about 200,000 messages a day. Analysis shows a long head of repeated questions: checking balance, activating data packs, porting numbers, eSIM activation — in several languages and many spellings.

The team starts carefully:

  1. Cache only answers to a curated list of 300 stable, non-personal FAQ intents.
  2. Scope by language, customer type (prepaid or postpaid) and circle, plus index and prompt version.
  3. Threshold 0.95, plus a rule: if the query contains a number (a price, a plan, a date), the number must match the cached question's.
  4. Log every hit; reviewers check a daily sample of near-threshold hits.

In the first review, they find "₹399 plan validity" answered from the cached "₹299 plan validity" entry — exactly the case the number rule was designed for; the rule had a bug with the ₹ symbol. After fixing it, a large share of FAQ traffic is served from the cache in tens of milliseconds, and the reviewed wrong-hit rate is close to zero. Only then do they widen the cache to more intents.

Follow-up questions to expect

  • "How do you choose the threshold?" — Label a few hundred real query pairs as same-answer or different-answer, and pick the threshold that keeps wrong hits near zero; accept a lower hit rate as the price.
  • "What about multi-turn conversations?" — Cache on the condensed standalone query, not the raw follow-up; "and for postpaid?" means different things in different conversations.
  • "How do you invalidate?" — Store the source document IDs with each cached answer; when a document changes, delete entries that cite it. Also bump the index or prompt version in the scope on each release.