RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

How do you optimize embedding costs?


What you need to know

Embedding cost has three parts, and people often look only at the first.

  1. API or GPU cost to turn text into vectors, charged per token.
  2. Storage and memory to hold the vectors. Vector indexes are fastest when they fit in RAM.
  3. Re-embedding everything when you change the model or chunking.

Worked numbers

A company indexes 50,000 documents averaging 2,000 tokens: 100 million tokens. Using an assumed price of $0.02 per million tokens for a small embedding model, one full embedding costs about $2. With a larger model at an assumed $0.13 per million, about $13. Cheap. But:

  • If the ingest job naively re-embeds everything nightly, that is 365 full runs a year.
  • The contextual retrieval trick (asking an LLM to write a short context line for each chunk before embedding) costs an LLM call per chunk, which is far more than the embedding.
  • Storage: 100 million tokens in 400-token chunks is 250,000 vectors. At 1,536 dimensions and 4 bytes per number, that is 250,000 × 1,536 × 4 ≈ 1.5 GB of raw vectors, before index overhead and replicas. At 256 dimensions it is about 0.26 GB.

So the real savings come from not repeating work and from smaller vectors.

The levers

  • Content-hash cache. Key = hash of (model name + chunk text). Unchanged text is never embedded twice.
  • Incremental ingestion. Only process documents whose updated_at or content hash changed.
  • Batching. Send a few hundred texts per request. Fewer requests means fewer rate-limit problems and retries. For large backfills, some providers offer an asynchronous batch API at a discount.
  • Remove boilerplate. Footers, disclaimers and navigation menus repeated on every page waste tokens and hurt retrieval.
  • Smaller model. Test two or three models on your own golden set. The leaderboard winner is often not better on a narrow domain.
  • Shorter vectors. Models trained with Matryoshka representation learning keep most of their quality when you keep only the first part of the vector (for example 1,536 cut to 512). Quantization stores each number in fewer bits: int8 is 4 times smaller than float32, binary is 32 times smaller, usually with a rescoring step to recover accuracy.
  • Self-host an open model on a GPU when volume is very high and steady, and you can run it reliably.

The cache code

Python
import hashlibMODEL = "text-embedding-3-small"cache: dict[str, list[float]] = {}          # use Redis or a table in productiondef embed_with_cache(texts, batch_size=256):    keys = [hashlib.sha256(f"{MODEL}|{t}".encode()).hexdigest() for t in texts]    todo = [(k, t) for k, t in zip(keys, texts) if k not in cache]    for i in range(0, len(todo), batch_size):        part = todo[i:i + batch_size]        vectors = embed_batch([t for _, t in part])      # your provider call        cache.update({k: v for (k, _), v in zip(part, vectors)})    return [cache[k] for k in keys]

The model name is part of the key, so switching models never returns an old vector. I tested this pattern with 1,000 chunks: the first run made 4 batch calls, and after editing one chunk the second run embedded exactly one text.

A real-life example

A hospital's guideline search re-embedded all 1,200 guidelines (about 40 million tokens) every night because "some might have changed". In a typical week, 5 to 10 guidelines change. After adding a content hash per document and a chunk-level cache, the nightly job embeds only the changed ones: roughly 0.5% of the previous volume. More useful than the money saved, the job now finishes in minutes instead of three hours, so updated guidelines are searchable the same morning.

Follow-up questions to expect

  • "Is query-time embedding a big cost?" — Usually not. A query is 10 to 30 tokens. Ingestion and storage dominate.
  • "When would you pay for a bigger embedding model?" — When your golden set shows a clear recall gain that a reranker or hybrid search does not already provide.
  • "What happens when you change the embedding model?" — Every vector must be rebuilt. Build a new index in parallel, evaluate, then switch.