Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Scenario – 4: Escalating Embedding Costs
Scenario: a Python ingestion script re-embeds the whole corpus on every run, and the embedding bill is climbing fast. How do you bring it down?
What you need to know
If 98% of a corpus is unchanged since yesterday, re-embedding it all spends 98% of the bill on vectors you already have. Embeddings are deterministic enough for this: the same model and the same text give you a vector you can safely reuse.
The cache
1import hashlib23def embed_all(texts: list[str], model="text-embedding-3-small") -> list[list[float]]:4 keys = [hashlib.sha256(f"{model}:{t}".encode()).hexdigest() for t in texts]5 found = dict(zip(keys, cache.mget(keys))) # None where missing6 todo = [(k, t) for k, t in zip(keys, texts) if found[k] is None]7 for batch in chunked(todo, 256): # batch, don't loop one by one8 vectors = api_embed([t for _, t in batch], model=model)9 for (k, _), v in zip(batch, vectors):10 cache.set(k, v)11 found[k] = v12 return [found[k] for k in keys]The model name is part of the key, so switching models never reuses a vector from the old one. Only new or changed texts reach the API, in batches of 256.
The other levers, by impact
| Lever | Why it saves | Watch out for |
|---|---|---|
| Deduplicate | Headers, footers, disclaimers repeat across thousands of documents | Near-duplicates need MinHash or similar, not just exact hashes |
| Batch requests | Fewer calls, less overhead, fewer rate-limit backoffs | Provider limits on batch size and tokens per request |
| Smaller model or fewer dimensions | Lower price per token and smaller vectors to store | Recall may drop; measure on your eval set |
| Self-host an open model | Fixed GPU cost instead of per-token fees | Only pays off with large, steady volume |
A monitoring detail that saves money
Track cache hit rate per run. If it suddenly falls from 97% to 3%, someone changed the chunker, the text normalisation or the model name, and every hash changed. An alert on that drop stops a surprise bill the same day.
- Cache — hash, look up, embed only misses.
- Deduplicate — exact hashes first, near-duplicates second.
- Batch — 100 to 256 texts per request.
- Right-size — test a smaller model or fewer dimensions against your retrieval eval.
- Monitor — hit rate, cost per million chunks, alert on sudden changes.
A real-life example
Scenario (illustrative numbers). A legal-research startup re-embeds 12 million chunks of court judgments every night because the script was written for a 5,000-document prototype. The embedding bill has grown to about $2,400 a month and the job takes 7 hours.
With the cache, a nightly run embeds only about 40,000 new chunks. Deduplication removes 1.8 million chunks of repeated headnote boilerplate from the index. Moving from a large model at full dimensions to a small model at 768 dimensions costs 1.5 points of recall@10 on their 300-query eval, which they accept. The monthly bill falls to about $60 and the job runs in 12 minutes. Two months later, the hit-rate alert fires after someone changes whitespace normalisation; they fix it before the next run.
Follow-up questions to expect
- "Where should the cache live?" — Anywhere durable and fast: Redis, a key-value table in Postgres, or even the vector store itself if it lets you look up by hash.
- "Can you reuse vectors after changing the model?" — No. Vectors from different models are not comparable; that is a full re-embed, done as a planned migration.
- "When does self-hosting pay off?" — When volume is large and steady enough to keep a GPU busy; below that, API pricing for small embedding models is very cheap.