Course Content
Advanced RAG
3 sections · 38 lessons
What is Cache-Augmented Generation, and when is it better than RAG?
What you need to know
How CAG works
- Put the entire corpus (say, 80,000 tokens of policies) into a fixed prefix.
- Run prefill once and keep the KV cache.
- For each question, append the question to the cached prefix and generate.
The model attends over every document for every question. Nothing can be "not retrieved", because everything is present.
Two ways to get the cache
- Self-hosted models — serving engines such as vLLM or SGLang keep prefix caches in GPU memory and reuse them automatically when prompts share a prefix.
- Hosted APIs — you cannot hold the KV cache yourself. Providers offer prompt caching instead: if the start of your prompt is byte-identical to a recent request, cached tokens are billed at a much lower rate and prefill is faster. The cache lives for minutes (some providers offer longer options), so it works best with steady traffic.
When CAG beats RAG
- The corpus fits comfortably in the window — a handbook, a product catalogue, a policy set. Not a wiki.
- It changes rarely, so you rebuild the cache rarely.
- Every user may see all of it.
- Questions often need facts from several places at once, which top-k retrieval can split apart.
When RAG stays better
- Size — millions of documents do not fit in any window.
- Freshness — one edited document means re-prefilling the whole prefix.
- Access control — a shared cache cannot hide half its content from one user. Per-user caches destroy the cost benefit.
- Citations — retrieval gives you chunk IDs for free. With CAG you rely on the model to cite exactly.
- Quality at length — models still get less reliable as context grows and fills with similar-looking passages.
Cost note: a cached token is cheaper, not free. Each query still pays attention over the full prefix, so a 100k-token prefix makes every question more expensive than a 3k-token RAG prompt.
A real-life example
An Indian telecom's support bot has two kinds of knowledge:
- The core — the current prepaid and postpaid plan catalogue plus the 150 most-asked FAQ answers. About 40,000 tokens, updated on the first of each month, identical for every customer.
- The long tail — 12,000 help articles, device guides and circle-specific notices, updated daily.
The team puts the core in a fixed prefix at the top of every prompt, so the provider's prompt cache serves it on almost every call. Questions like "Which plan gives the most data under ₹300?" need comparisons across the whole catalogue, and with the full catalogue in context the model compares every plan instead of the three a retriever happened to return. Everything else goes through normal hybrid retrieval, appended after the cached prefix.
The one rule they enforce: nothing customer-specific (balance, usage, KYC status) goes into the shared prefix. Personal data comes from a tool call, after the cache boundary.
Follow-up questions to expect
- "Why keep the static part first in the prompt?" — Prompt caches match on the prefix. Anything that varies — a timestamp, a user name — placed before the static block breaks the match and every call pays full price.
- "How would you evaluate CAG against RAG?" — Run both on the same question set and compare accuracy, faithfulness, p95 latency and cost per query. Include questions that need facts from several documents, where CAG should win, and a size stress test, where it may not.
- "Is CAG the same as long-context RAG?" — Close. CAG is "put everything in and cache it". Long-context RAG still retrieves, but retrieves generously and lets the big window absorb the imprecision.