Advanced RAG

Course Content

Advanced RAG

3 sections · 38 lessons

What is token budget management, and how do you balance context vs generation?


What you need to know

The buckets

Everything the model reads and writes shares one window. A typical plan for a 128k-token model might look like this:

Text
window 128k  reserve   4k  completion (max output tokens)            1k  system prompt          (never truncated)            2k  few-shot examples      (drop the last one first)  history   up to 8k                   (summarise older turns)  retrieved remainder, capped ~20k     (drop lowest-ranked first)

The cap on retrieved context is deliberate. The window could take 110k tokens of documents, but that would cost far more per call and usually answer worse.

A packing function

Python
def pack(window, reserve_output, fixed, history, chunks, count):    """fixed: never cut. history: newest first. chunks: best-ranked first."""    budget = window - reserve_output - count(fixed)    hist_cap = budget // 4                      # history gets at most 25%    kept_hist, kept_chunks = [], []    for turn in history:        if count(turn) > hist_cap: break        kept_hist.append(turn); hist_cap -= count(turn); budget -= count(turn)    for c in chunks:        if count(c) > budget: break             # never cut a chunk in half        kept_chunks.append(c); budget -= count(c)    return kept_hist, kept_chunks, budgetcount = lambda s: len(s.split())                # stand-in for a real tokenizerfixed = "rule " * 1200history = ["turn " * 400] * 6chunks = ["chunk " * 900] * 10h, c, spare = pack(8000, 1500, fixed, history, chunks, count)print(len(h), "turns,", len(c), "chunks,", spare, "tokens spare")   # 3 turns, 4 chunks, 500 tokens spare

The order of operations is the policy: reserve output, protect fixed instructions, cap history, then fill with the best chunks. In production, count must be the model's real tokenizer or the provider's token-counting endpoint; words or len(text) / 4 are fine as rough guards and wrong exactly near the limit.

How much retrieved context?

Accuracy usually rises quickly for the first few chunks, then flattens, and can fall as distractors pile up. Cost and prefill latency keep rising in a straight line. Find the elbow empirically: run your evaluation set at top-k = 2, 4, 6, 8, 12, 20 and plot accuracy and cost. The best value is often lower than teams expect.

Degrade gracefully, in a fixed order

  1. Drop the lowest-ranked chunks.
  2. Compress the remaining chunks (extractive filtering).
  3. Summarise older history.
  4. Drop the last few-shot example.

Never cut a document mid-sentence — half a condition is worse than no condition. And reasoning models need extra output room: their internal thinking often counts against the output budget.

A real-life example

A bank's compliance assistant uses a model with a 128k window. Under load, officers start receiving answers that stop mid-sentence. Logs show why: long conversations plus 30 retrieved chunks leave only a few hundred tokens for the answer, and the model hits the output limit.

The team introduces explicit budgets: 4,000 output tokens reserved (the model is a reasoning model, so it needs room to think), system prompt and citation rules protected, history capped at 6,000 tokens with older turns summarised, and retrieved context capped at 16,000 tokens.

Then they sweep top-k on 300 labelled questions. Accuracy at top-6 is the same as at top-30, and slightly better on questions where older policy versions used to confuse the model. Cost per question falls by more than half, p95 latency drops noticeably because prefill is shorter, and truncated answers disappear.

Follow-up questions to expect

  • "What if the one crucial chunk doesn't fit?" — Compress other chunks first, or retrieve smaller units. If a single source is larger than the budget, extract only its relevant sections.
  • "Should history or retrieved context win?" — It depends on the turn. For a follow-up like "explain that more simply", history wins; for a new factual question, retrieved context wins. A cap on each keeps either from starving the other.
  • "Why not just use the biggest window available?" — Cost and latency scale with input tokens, and quality does not keep rising. A bigger window is room for error, not a target.