Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Scenario – 5: Rising Embedding Spend


What you need to know

The scenario: the monthly embedding bill keeps climbing even though the document count is growing slowly.

Where the money goes

BucketSignFix
Re-embedding unchanged contentNightly job embeds the whole corpusContent hash per chunk; embed only changes
Query embeddingsCost scales with trafficCache by normalised query; smaller model if recall holds
Backfills at full priceBig jumps after migrationsBatch API, usually about half price
Junk contentNavigation, footers, tiny chunksStrip at parse time; drop chunks under ~50 tokens

Incremental indexing

Python
import hashlibdef sync(chunks, store):    to_embed = []    for c in chunks:        h = hashlib.sha256(" ".join(c.text.split()).lower().encode()).hexdigest()        if store.get_hash(c.id) != h:            to_embed.append((c, h))    for batch in batched(to_embed, 256):        vectors = embed([c.text for c, _ in batch])        store.upsert([(c.id, v, h) for (c, h), v in zip(batch, vectors)])    store.delete_missing(keep_ids={c.id for c in chunks})   # remove deleted chunks too    return len(to_embed)

Changing chunk boundaries changes every hash, so fix chunking once and keep it stable.

Right-size the model

Smaller embedding models are often close to the largest ones on a specific domain. Measure recall@10 on your golden set before paying for the biggest model. Models trained Matryoshka-style let you keep only the first part of each vector, which also cuts vector-database memory. At steady high volume, self-hosting an open embedding model on one GPU can be cheaper than an API; do the arithmetic with your own traffic.

Guard quality

Every cost change must pass the recall gate. Track embedding tokens per day, cost per 1,000 documents indexed, cache hit rate and recall@10 on the same dashboard.

A real-life example

Scenario, numbers made up. A SaaS company's help-centre RAG spends about $3,000 a month on embeddings. The docs change by roughly 2% a day, but a nightly job re-embeds all 1.5M chunks.

Adding content hashes cuts the nightly job to about 30,000 chunks. A query cache in Redis reaches a 35% hit rate. Testing a smaller embedding model on their 400-query golden set shows recall@10 drops by less than one point, so they switch for queries and documents together with a planned re-index. Monthly spend falls to under $300, and recall stays within the gate.

Follow-up questions to expect

  • "Can you use a small model for queries and a big one for documents?" — Not across different models; queries and documents must be embedded by the same model (or a pair explicitly trained to work together).
  • "How do you cache queries safely?" — Key by the normalised query text plus the model version, and clear the cache when the model changes.
  • "When is self-hosting worth it?" — When volume is high and steady enough that a GPU stays busy; for spiky or low traffic, APIs are usually cheaper.