Course Content
Advanced RAG
3 sections · 38 lessons
What are the biggest cost drivers in RAG systems, and how can they be reduced?
What you need to know
Where the money goes
Work it out per query, then multiply by traffic. A typical RAG call has:
system prompt + rules + examples 1,500 tokens (same every call)conversation history 500 tokensretrieved context (8 chunks x 400) 3,200 tokensquestion 50 tokens-------------------------------------------------input 5,250 tokensoutput 250 tokens+ query rewrite call (small model) ~400 tokens in, 40 out+ reranker over 50 candidates+ 1 query embeddingRetrieved context is about 60% of the input here. Output tokens usually cost several times more per token than input, but there are far fewer of them in RAG answers. So for most RAG systems input tokens dominate.
The cost drivers and their fixes
| Driver | Why it costs | Main fix |
|---|---|---|
| Retrieved context | Most of every prompt | Rerank, send top 3–5, extractive compression |
| Static prefix | Re-sent on every call | Prompt caching; keep it first and byte-identical |
| Auxiliary LLM calls | Rewriting, multi-query, grading, self-checks | Small models; skip steps via routing |
| Repeated questions | Full pipeline for the same FAQ | Semantic cache (scoped) |
| Output tokens | Priced higher per token | Ask for concise answers; set output limits |
| Reasoning tokens | Thinking models may use many hidden tokens | Use lower effort, or a non-reasoning model for simple questions |
| Vector storage | RAM for large HNSW indexes | Scalar or binary quantisation; shorter (Matryoshka) embeddings |
| Reranking | Linear in candidates and tokens | Pick N at the recall knee; cascade |
| Re-embedding | Whole corpus on a model change | Plan migrations; incremental updates by content hash |
Prompt caching done right
Providers bill cached prefix tokens at a much lower rate, but only when the beginning of the prompt is exactly the same as a recent request. Put stable parts first (system prompt, tool definitions, examples), variable parts last (retrieved context, question). A timestamp or user name at the top breaks the cache for every call.
Measure before optimising
Attach cost to each stage in your traces: tokens per call, calls per query, model per call. Teams are often surprised — a "cheap" grading step running on a frontier model, or a self-check loop that runs three times on hard queries, can outweigh the main answer call.
A real-life example
An Indian telecom's support bot handles 200,000 messages a day. The finance team flags the monthly model bill. The engineers trace a sample of 10,000 requests and find:
- The main answer call uses a large model with 8 retrieved chunks. Retrieved context is most of its input.
- The system prompt (1,500 tokens) sits after a line containing the customer's name and the current time, so the prompt cache never hits.
- A "safety self-check" step re-reads every answer with the same large model.
- About a third of messages are repeat FAQ intents.
Changes, in order of payoff:
- Move per-user lines below the static prefix → the prefix is now cached on almost every call.
- Rerank to top 4 and add extractive compression → context shrinks by about three-quarters; accuracy on the 600-question eval set is unchanged.
- Run the query rewrite and the self-check on a small model → auxiliary cost drops sharply; eval shows no loss.
- Add a scoped semantic cache for 300 stable FAQ intents → those requests skip the model entirely.
The bill falls to a fraction of its original level, and median latency improves because prompts are shorter. Nothing was done to the vector database, which turned out to be a small part of the cost.
Follow-up questions to expect
- "Would you switch to a smaller model for everything?" — Route instead: simple FAQ-type questions to a small model, complex or high-risk ones to a large model. Measure quality per route.
- "Does a long-context model save money by removing retrieval?" — Usually the opposite: sending a whole corpus per query costs far more than a few retrieved chunks, even with caching.
- "How does quantisation affect quality?" — It lowers recall slightly; a common pattern is to search on quantised vectors and rescore the top candidates with full-precision vectors.