Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Scenario – 5: Rising Embedding Spend
What you need to know
The scenario: the monthly embedding bill keeps climbing even though the document count is growing slowly.
Where the money goes
| Bucket | Sign | Fix |
|---|---|---|
| Re-embedding unchanged content | Nightly job embeds the whole corpus | Content hash per chunk; embed only changes |
| Query embeddings | Cost scales with traffic | Cache by normalised query; smaller model if recall holds |
| Backfills at full price | Big jumps after migrations | Batch API, usually about half price |
| Junk content | Navigation, footers, tiny chunks | Strip at parse time; drop chunks under ~50 tokens |
Incremental indexing
1import hashlib23def sync(chunks, store):4 to_embed = []5 for c in chunks:6 h = hashlib.sha256(" ".join(c.text.split()).lower().encode()).hexdigest()7 if store.get_hash(c.id) != h:8 to_embed.append((c, h))9 for batch in batched(to_embed, 256):10 vectors = embed([c.text for c, _ in batch])11 store.upsert([(c.id, v, h) for (c, h), v in zip(batch, vectors)])12 store.delete_missing(keep_ids={c.id for c in chunks}) # remove deleted chunks too13 return len(to_embed)Changing chunk boundaries changes every hash, so fix chunking once and keep it stable.
Right-size the model
Smaller embedding models are often close to the largest ones on a specific domain. Measure recall@10 on your golden set before paying for the biggest model. Models trained Matryoshka-style let you keep only the first part of each vector, which also cuts vector-database memory. At steady high volume, self-hosting an open embedding model on one GPU can be cheaper than an API; do the arithmetic with your own traffic.
Guard quality
Every cost change must pass the recall gate. Track embedding tokens per day, cost per 1,000 documents indexed, cache hit rate and recall@10 on the same dashboard.
A real-life example
Scenario, numbers made up. A SaaS company's help-centre RAG spends about $3,000 a month on embeddings. The docs change by roughly 2% a day, but a nightly job re-embeds all 1.5M chunks.
Adding content hashes cuts the nightly job to about 30,000 chunks. A query cache in Redis reaches a 35% hit rate. Testing a smaller embedding model on their 400-query golden set shows recall@10 drops by less than one point, so they switch for queries and documents together with a planned re-index. Monthly spend falls to under $300, and recall stays within the gate.
Follow-up questions to expect
- "Can you use a small model for queries and a big one for documents?" — Not across different models; queries and documents must be embedded by the same model (or a pair explicitly trained to work together).
- "How do you cache queries safely?" — Key by the normalised query text plus the model version, and clear the cache when the model changes.
- "When is self-hosting worth it?" — When volume is high and steady enough that a GPU stays busy; for spiky or low traffic, APIs are usually cheaper.