Advanced RAG

Course Content

Advanced RAG

3 sections · 38 lessons

What are the biggest cost drivers in RAG systems, and how can they be reduced?


What you need to know

Where the money goes

Work it out per query, then multiply by traffic. A typical RAG call has:

Text
system prompt + rules + examples     1,500 tokens   (same every call)conversation history                   500 tokensretrieved context (8 chunks x 400)   3,200 tokensquestion                                50 tokens-------------------------------------------------input                                5,250 tokensoutput                                 250 tokens+ query rewrite call (small model)     ~400 tokens in, 40 out+ reranker over 50 candidates+ 1 query embedding

Retrieved context is about 60% of the input here. Output tokens usually cost several times more per token than input, but there are far fewer of them in RAG answers. So for most RAG systems input tokens dominate.

The cost drivers and their fixes

DriverWhy it costsMain fix
Retrieved contextMost of every promptRerank, send top 3–5, extractive compression
Static prefixRe-sent on every callPrompt caching; keep it first and byte-identical
Auxiliary LLM callsRewriting, multi-query, grading, self-checksSmall models; skip steps via routing
Repeated questionsFull pipeline for the same FAQSemantic cache (scoped)
Output tokensPriced higher per tokenAsk for concise answers; set output limits
Reasoning tokensThinking models may use many hidden tokensUse lower effort, or a non-reasoning model for simple questions
Vector storageRAM for large HNSW indexesScalar or binary quantisation; shorter (Matryoshka) embeddings
RerankingLinear in candidates and tokensPick N at the recall knee; cascade
Re-embeddingWhole corpus on a model changePlan migrations; incremental updates by content hash

Prompt caching done right

Providers bill cached prefix tokens at a much lower rate, but only when the beginning of the prompt is exactly the same as a recent request. Put stable parts first (system prompt, tool definitions, examples), variable parts last (retrieved context, question). A timestamp or user name at the top breaks the cache for every call.

Measure before optimising

Attach cost to each stage in your traces: tokens per call, calls per query, model per call. Teams are often surprised — a "cheap" grading step running on a frontier model, or a self-check loop that runs three times on hard queries, can outweigh the main answer call.

A real-life example

An Indian telecom's support bot handles 200,000 messages a day. The finance team flags the monthly model bill. The engineers trace a sample of 10,000 requests and find:

  • The main answer call uses a large model with 8 retrieved chunks. Retrieved context is most of its input.
  • The system prompt (1,500 tokens) sits after a line containing the customer's name and the current time, so the prompt cache never hits.
  • A "safety self-check" step re-reads every answer with the same large model.
  • About a third of messages are repeat FAQ intents.

Changes, in order of payoff:

  1. Move per-user lines below the static prefix → the prefix is now cached on almost every call.
  2. Rerank to top 4 and add extractive compression → context shrinks by about three-quarters; accuracy on the 600-question eval set is unchanged.
  3. Run the query rewrite and the self-check on a small model → auxiliary cost drops sharply; eval shows no loss.
  4. Add a scoped semantic cache for 300 stable FAQ intents → those requests skip the model entirely.

The bill falls to a fraction of its original level, and median latency improves because prompts are shorter. Nothing was done to the vector database, which turned out to be a small part of the cost.

Follow-up questions to expect

  • "Would you switch to a smaller model for everything?" — Route instead: simple FAQ-type questions to a small model, complex or high-risk ones to a large model. Measure quality per route.
  • "Does a long-context model save money by removing retrieval?" — Usually the opposite: sending a whole corpus per query costs far more than a few retrieved chunks, even with caching.
  • "How does quantisation affect quality?" — It lowers recall slightly; a common pattern is to search on quantised vectors and rescore the top candidates with full-precision vectors.