Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Scenario – 4: Escalating Embedding Costs


Scenario: a Python ingestion script re-embeds the whole corpus on every run, and the embedding bill is climbing fast. How do you bring it down?

Embed only what changedChunk textHash of modelname plus textHit: reuse thestored vectorMiss: batch of256 to the APIStore the vectorunder its hashAlert when the hit rate suddenly drops.
A nightly run over 12 million chunks shrinks to the 40,000 that are actually new.

What you need to know

If 98% of a corpus is unchanged since yesterday, re-embedding it all spends 98% of the bill on vectors you already have. Embeddings are deterministic enough for this: the same model and the same text give you a vector you can safely reuse.

The cache

Python
import hashlibdef embed_all(texts: list[str], model="text-embedding-3-small") -> list[list[float]]:    keys = [hashlib.sha256(f"{model}:{t}".encode()).hexdigest() for t in texts]    found = dict(zip(keys, cache.mget(keys)))              # None where missing    todo = [(k, t) for k, t in zip(keys, texts) if found[k] is None]    for batch in chunked(todo, 256):                       # batch, don't loop one by one        vectors = api_embed([t for _, t in batch], model=model)        for (k, _), v in zip(batch, vectors):            cache.set(k, v)            found[k] = v    return [found[k] for k in keys]

The model name is part of the key, so switching models never reuses a vector from the old one. Only new or changed texts reach the API, in batches of 256.

The other levers, by impact

LeverWhy it savesWatch out for
DeduplicateHeaders, footers, disclaimers repeat across thousands of documentsNear-duplicates need MinHash or similar, not just exact hashes
Batch requestsFewer calls, less overhead, fewer rate-limit backoffsProvider limits on batch size and tokens per request
Smaller model or fewer dimensionsLower price per token and smaller vectors to storeRecall may drop; measure on your eval set
Self-host an open modelFixed GPU cost instead of per-token feesOnly pays off with large, steady volume

A monitoring detail that saves money

Track cache hit rate per run. If it suddenly falls from 97% to 3%, someone changed the chunker, the text normalisation or the model name, and every hash changed. An alert on that drop stops a surprise bill the same day.

  1. Cache — hash, look up, embed only misses.
  2. Deduplicate — exact hashes first, near-duplicates second.
  3. Batch — 100 to 256 texts per request.
  4. Right-size — test a smaller model or fewer dimensions against your retrieval eval.
  5. Monitor — hit rate, cost per million chunks, alert on sudden changes.

A real-life example

Scenario (illustrative numbers). A legal-research startup re-embeds 12 million chunks of court judgments every night because the script was written for a 5,000-document prototype. The embedding bill has grown to about $2,400 a month and the job takes 7 hours.

With the cache, a nightly run embeds only about 40,000 new chunks. Deduplication removes 1.8 million chunks of repeated headnote boilerplate from the index. Moving from a large model at full dimensions to a small model at 768 dimensions costs 1.5 points of recall@10 on their 300-query eval, which they accept. The monthly bill falls to about $60 and the job runs in 12 minutes. Two months later, the hit-rate alert fires after someone changes whitespace normalisation; they fix it before the next run.

Follow-up questions to expect

  • "Where should the cache live?" — Anywhere durable and fast: Redis, a key-value table in Postgres, or even the vector store itself if it lets you look up by hash.
  • "Can you reuse vectors after changing the model?" — No. Vectors from different models are not comparable; that is a full re-embed, done as a planned migration.
  • "When does self-hosting pay off?" — When volume is large and steady enough to keep a GPU busy; below that, API pricing for small embedding models is very cheap.