Course Content
Advanced RAG
3 sections · 38 lessons
How do you maximize use of the context window when data is limited?
What you need to know
Reframe the problem
"Limited data" usually means a small corpus — tens or hundreds of documents — and a context window far larger than the relevant text. The question is not "how do I fill the window?" but "what extra tokens actually improve the answer?"
Options, roughly in order of value
- Consider no retrieval at all — if the entire corpus is, say, 60,000 tokens and changes weekly, put it all in the prompt, static part first, with prompt caching. No retrieval misses, no chunking decisions.
- Return larger units — if you still retrieve, return whole sections or documents instead of 300-token chunks. Small chunks exist to save space; with space to spare, context is worth more.
- Add contextual headers — prepend the path to every chunk:
Deploy Guide v3 > Rollbacks > Database migrations (updated 2026-05-14). It helps the model know what it is reading. - Include metadata — version, date, status (draft or final), owner. When two passages disagree, the model can only prefer the current one if it can see which is current.
- Add worked examples — two or three examples of an excellent answer in your required format, and a short list of known gotchas from evaluation failures ("the staging and production URLs differ; always state which").
- Order deliberately — best material first, the question repeated after the context; avoid burying the key passage in the middle.
What not to do
- Don't pad. Filling the window with near-duplicate chunks or loosely related pages makes answers worse, not better. Deduplicate near-identical chunks (very high cosine similarity) before building the prompt.
- Don't raise top-k blindly. With a small corpus, top-20 may include most of the corpus's unrelated documents.
- Don't forget cost. A large prompt on every call costs more per query even if cached; keep only what pays for itself on your evals.
When "limited data" means few answers exist
Sometimes the corpus simply doesn't cover many questions. Then the best use of the window is a clear instruction: "If the documents do not contain the answer, say so and suggest who to ask." Log those questions — they are the list of documents to write next.
A real-life example
A 40-person startup builds an engineering-wiki assistant. The whole wiki is 120 pages, about 180,000 tokens. The first version copies a large-company design: 300-token chunks, top-5 retrieval. Engineers complain that answers about deployment miss steps and mix up staging and production.
The team rebuilds it for a small corpus:
- The 12 most-used pages — deploy guide, on-call handbook, service map, about 45,000 tokens — go into a cached prompt prefix on every call.
- For the rest, retrieval returns whole pages (top 3) rather than chunks, each with a header showing title, section path and last-updated date.
- The prompt includes two example answers in the house format (steps plus the exact command), and a "known traps" list of five items taken from the last month's thumbs-down answers — for example, "the staging cluster name ends in
-stg; never give production commands for staging questions". - An instruction handles conflicts: "If two pages disagree, prefer the most recently updated and mention the conflict."
On their 120-question evaluation set, the missing-steps problem mostly disappears, and the staging/production mix-up stops. The cost per question rises, but at a few hundred questions a day it is a small line in the budget.
Follow-up questions to expect
- "How do you decide between whole-corpus-in-prompt and retrieval?" — Measure both on the same eval set. If the corpus fits comfortably, rarely changes and has no per-user restrictions, whole-corpus often wins on accuracy and simplicity.
- "Where should the examples go — before or after the documents?" — Keep fixed content (instructions, examples) at the start so it can be cached; put the variable retrieved context and the question after it.
- "What if documents conflict and dates are missing?" — Fix it at the source: add version metadata at ingestion. The model cannot reliably guess which document is current.