Course Content
Building AI Features in Python Backends
5 sections · 23 lessons
Caching: exact, normalised and semantic
Caching is the cheapest speed-up in backend engineering, and with model calls it also cuts the bill. A cache hit costs about a millisecond and nothing else; a classifier call costs about 900 ms and $0.003. So the question is not whether caching helps, but what you can cache without serving a wrong answer.
That question is sharper than usual here. A stale product price in a cache is annoying. A cached label that says "reschedule" for a message that actually says "do not reschedule" sends a parcel to a customer who is not home, and nobody sees the cache was involved. The rule for this lesson: cache only when the same key must give the same correct answer.
What is safe to cache
| Result | Cache it? | Why |
|---|---|---|
| Classification | Yes | The label depends only on the text and the prompt. Same text, same correct label. |
| Extraction | No | "Tomorrow" depends on today's date, and tracking IDs differ in every message. |
| Draft reply | No | It includes the extracted fields, and repeating a reply word for word looks robotic. |
| Status question | Not needed | The regex already answers it for free. |
So ShipFast caches exactly one thing: the validated classification. It caches the Pydantic object after validation, never the raw text, and never an unknown label, because unknown often comes from a failure and a failure must not stick for hours.
Exact and normalised keys
An exact cache uses the message text as the key. For ShipFast, exact repeats are rare, about 3% of messages, mostly templated messages from a few business customers' systems.
A normalised key removes differences that cannot change the label: capital letters, punctuation, emoji, extra spaces and, for classification only, the specific tracking ID.
1# shipfast/cache.py2import hashlib3import re4import time567def normalise(text: str) -> str:8 text = text.lower().strip()9 text = re.sub(r"\bsf\d{8}\b", "<tracking>", text)10 text = re.sub(r"[^\w\s<>]", " ", text) # drop punctuation and emoji11 return re.sub(r"\s+", " ", text).strip()121314def cache_key(feature: str, prompt_version: str, model: str, text: str) -> str:15 raw = f"{feature}|{prompt_version}|{model}|{normalise(text)}"16 return hashlib.sha256(raw.encode()).hexdigest()171819class TTLCache:20 """In-process stand-in for Redis: GET, and SET with an expiry."""2122 def __init__(self, ttl_s: float = 3600, max_items: int = 50_000) -> None:23 self.ttl_s, self.max_items = ttl_s, max_items24 self._data: dict[str, tuple[float, str]] = {}2526 def get(self, key: str) -> str | None:27 item = self._data.get(key)28 if item is None or item[0] < time.monotonic():29 self._data.pop(key, None)30 return None31 return item[1]3233 def set(self, key: str, value: str) -> None:34 if len(self._data) >= self.max_items:35 self._data.pop(next(iter(self._data))) # evict the oldest insert36 self._data[key] = (time.monotonic() + self.ttl_s, value)"Please deliver TOMORROW after 6!!" and "please deliver tomorrow after 6" now share a key, and so do "SF12345678 not home today" and "SF99990000 not home today". Normalisation lifted ShipFast's hit rate from 3% to about 10%.
Look at what is in the key besides the text. The prompt version and the model are part of it, so when you ship classify-v4 or change model, old entries are simply never read again. You never need to flush the cache on deploy, and you can never serve a label produced by a prompt you have since fixed. The key is a hash, so raw customer text is not stored as a key in Redis, which matters for the logging rules in Section 5.
In production, TTLCache is replaced by Redis with the same two operations (GET and SET key value EX 21600), so every instance shares hits. ShipFast uses a 6-hour expiry: long enough to catch the same templated messages through a working day, short enough that a problem does not persist.
The classifier gains a few lines:
1# shipfast/classify.py (updated)2from shipfast.cache import TTLCache, cache_key345async def classify(llm: LLM, text: str, *, budget: RequestBudget | None = None,6 cache: TTLCache | None = None) -> Classification:7 key = cache_key("classify", PROMPT_VERSION, llm.model, text)8 if cache and (hit := cache.get(key)):9 return Classification.model_validate_json(hit)10 try:11 out = await call_structured(12 llm, feature="classify", system=SYSTEM, user_text=f"<message>\n{text}\n</message>",13 model_cls=Classification, max_tokens=200, budget=budget)14 except StructuredOutputError:15 return Classification(reason="no valid label after repair", intent=Intent.UNKNOWN)16 if cache and out.value.intent is not Intent.UNKNOWN:17 cache.set(key, out.value.model_dump_json())18 return out.valueA hit is validated again on the way out. It costs microseconds, and it means a corrupted or hand-edited Redis entry cannot put an invalid label into the pipeline.
Semantic caching, and why ShipFast does not use it for labels
A semantic cache turns each message into an embedding (a list of numbers representing its meaning), and on a new message looks for a cached one whose embedding is close enough, say cosine similarity above 0.92. If found, it reuses that answer. Hit rates can be much higher, because "pls deliver tomorrow evening" and "can you bring it tmrw evening" match.
ShipFast measured it on a week of traffic, with the results checked against agents' final labels:
| Cache type | Hit rate | Wrong labels served from cache |
|---|---|---|
| Exact | 3.1% | 0 |
| Normalised | 9.8% | 0 |
| Semantic, threshold 0.92 | 27% | 1.6% of hits |
| Semantic, threshold 0.97 | 12% | 0.3% of hits |
The wrong hits were the dangerous kind. "Deliver tomorrow" and "do not deliver tomorrow" are very close in embedding space, because they share almost all their words. So are "parcel damaged" and "parcel not damaged". At 0.97, the semantic cache barely beats normalisation and still serves some wrong labels. For a classifier that decides where a customer's request goes, that trade is not worth it.
Semantic caching can make sense where a near-miss answer is still acceptable, such as returning a help-centre article for a general question. When a small difference in words means a different action, keep to exact or normalised keys.
Check your understanding
0 of 3 answered
1.Why does ShipFast not cache extraction results?
2.You ship classify-v4. What happens to cache entries created by classify-v3?
3.A semantic cache with threshold 0.92 matches "do not deliver tomorrow" to a cached "deliver tomorrow". What does this show?