Building AI Features in Python Backends

Caching: exact, normalised and semantic


Caching is the cheapest speed-up in backend engineering, and with model calls it also cuts the bill. A cache hit costs about a millisecond and nothing else; a classifier call costs about 900 ms and $0.003. So the question is not whether caching helps, but what you can cache without serving a wrong answer.

That question is sharper than usual here. A stale product price in a cache is annoying. A cached label that says "reschedule" for a message that actually says "do not reschedule" sends a parcel to a customer who is not home, and nobody sees the cache was involved. The rule for this lesson: cache only when the same key must give the same correct answer.

A week of ShipFast labels through four caches3.1%09.8%027%1.6% of hits12%0.3% of hitsHit rateWrong labels servedExactNormalisedSemantic at 0.92Semantic at 0.97
'deliver tomorrow' and 'do not deliver tomorrow' are neighbours in embedding space, so labels use normalised keys only.

What is safe to cache

ResultCache it?Why
ClassificationYesThe label depends only on the text and the prompt. Same text, same correct label.
ExtractionNo"Tomorrow" depends on today's date, and tracking IDs differ in every message.
Draft replyNoIt includes the extracted fields, and repeating a reply word for word looks robotic.
Status questionNot neededThe regex already answers it for free.

So ShipFast caches exactly one thing: the validated classification. It caches the Pydantic object after validation, never the raw text, and never an unknown label, because unknown often comes from a failure and a failure must not stick for hours.

Exact and normalised keys

An exact cache uses the message text as the key. For ShipFast, exact repeats are rare, about 3% of messages, mostly templated messages from a few business customers' systems.

A normalised key removes differences that cannot change the label: capital letters, punctuation, emoji, extra spaces and, for classification only, the specific tracking ID.

Python
# shipfast/cache.pyimport hashlibimport reimport timedef normalise(text: str) -> str:    text = text.lower().strip()    text = re.sub(r"\bsf\d{8}\b", "<tracking>", text)    text = re.sub(r"[^\w\s<>]", " ", text)   # drop punctuation and emoji    return re.sub(r"\s+", " ", text).strip()def cache_key(feature: str, prompt_version: str, model: str, text: str) -> str:    raw = f"{feature}|{prompt_version}|{model}|{normalise(text)}"    return hashlib.sha256(raw.encode()).hexdigest()class TTLCache:    """In-process stand-in for Redis: GET, and SET with an expiry."""    def __init__(self, ttl_s: float = 3600, max_items: int = 50_000) -> None:        self.ttl_s, self.max_items = ttl_s, max_items        self._data: dict[str, tuple[float, str]] = {}    def get(self, key: str) -> str | None:        item = self._data.get(key)        if item is None or item[0] < time.monotonic():            self._data.pop(key, None)            return None        return item[1]    def set(self, key: str, value: str) -> None:        if len(self._data) >= self.max_items:            self._data.pop(next(iter(self._data)))   # evict the oldest insert        self._data[key] = (time.monotonic() + self.ttl_s, value)

"Please deliver TOMORROW after 6!!" and "please deliver tomorrow after 6" now share a key, and so do "SF12345678 not home today" and "SF99990000 not home today". Normalisation lifted ShipFast's hit rate from 3% to about 10%.

Look at what is in the key besides the text. The prompt version and the model are part of it, so when you ship classify-v4 or change model, old entries are simply never read again. You never need to flush the cache on deploy, and you can never serve a label produced by a prompt you have since fixed. The key is a hash, so raw customer text is not stored as a key in Redis, which matters for the logging rules in Section 5.

In production, TTLCache is replaced by Redis with the same two operations (GET and SET key value EX 21600), so every instance shares hits. ShipFast uses a 6-hour expiry: long enough to catch the same templated messages through a working day, short enough that a problem does not persist.

The classifier gains a few lines:

Python
# shipfast/classify.py (updated)from shipfast.cache import TTLCache, cache_keyasync def classify(llm: LLM, text: str, *, budget: RequestBudget | None = None,                   cache: TTLCache | None = None) -> Classification:    key = cache_key("classify", PROMPT_VERSION, llm.model, text)    if cache and (hit := cache.get(key)):        return Classification.model_validate_json(hit)    try:        out = await call_structured(            llm, feature="classify", system=SYSTEM, user_text=f"<message>\n{text}\n</message>",            model_cls=Classification, max_tokens=200, budget=budget)    except StructuredOutputError:        return Classification(reason="no valid label after repair", intent=Intent.UNKNOWN)    if cache and out.value.intent is not Intent.UNKNOWN:        cache.set(key, out.value.model_dump_json())    return out.value

A hit is validated again on the way out. It costs microseconds, and it means a corrupted or hand-edited Redis entry cannot put an invalid label into the pipeline.

Semantic caching, and why ShipFast does not use it for labels

A semantic cache turns each message into an embedding (a list of numbers representing its meaning), and on a new message looks for a cached one whose embedding is close enough, say cosine similarity above 0.92. If found, it reuses that answer. Hit rates can be much higher, because "pls deliver tomorrow evening" and "can you bring it tmrw evening" match.

ShipFast measured it on a week of traffic, with the results checked against agents' final labels:

Cache typeHit rateWrong labels served from cache
Exact3.1%0
Normalised9.8%0
Semantic, threshold 0.9227%1.6% of hits
Semantic, threshold 0.9712%0.3% of hits

The wrong hits were the dangerous kind. "Deliver tomorrow" and "do not deliver tomorrow" are very close in embedding space, because they share almost all their words. So are "parcel damaged" and "parcel not damaged". At 0.97, the semantic cache barely beats normalisation and still serves some wrong labels. For a classifier that decides where a customer's request goes, that trade is not worth it.

Semantic caching can make sense where a near-miss answer is still acceptable, such as returning a help-centre article for a general question. When a small difference in words means a different action, keep to exact or normalised keys.

Check your understanding

0 of 3 answered

1.Why does ShipFast not cache extraction results?

2.You ship classify-v4. What happens to cache entries created by classify-v3?

3.A semantic cache with threshold 0.92 matches "do not deliver tomorrow" to a cached "deliver tomorrow". What does this show?