Course Content
RAG Systems
12 sections · 66 lessons
What does the k parameter control in vector retrieval?
What you need to know
k changes recall directly
If the right chunk is at rank 7, then recall@5 is 0 and recall@10 is 1 for that question. Across a labelled set, recall@k always rises as k rises, but with diminishing returns. A typical curve rises steeply to about k = 10, then flattens.
k changes cost directly
With 400-token chunks:
| k | Context tokens | Effect |
|---|---|---|
| 3 | 1,200 | Cheap, fast; misses split answers |
| 5 | 2,000 | Common default for direct questions |
| 10 | 4,000 | Better for comparisons and multi-part questions |
| 30 | 12,000 | Too much to send directly; fine as reranker input |
At 10,000 questions a day, going from k = 5 to k = 30 adds 100 million input tokens a day. And more context is not always better for quality: irrelevant chunks distract the model, and facts in the middle of long prompts are used less reliably.
Two ks, not one
- Retrieve wide — hybrid search returns 30 candidates. This step is about recall: is the answer anywhere in the list?
- Rerank — a cross-encoder scores all 30 against the question.
- Send narrow — the top 3 to 5 go into the prompt. This step is about precision.
Measure recall@30 for the first step and recall@5 after reranking. If recall@30 is low, fix retrieval (chunking, hybrid, filters). If recall@30 is high but recall@5 is low, a reranker will help.
Adaptive k
You do not always need a fixed count. Keep chunks until the score drops sharply or below a calibrated threshold, with a minimum and maximum. A simple factual question may use 2 chunks; a comparison may use 8.
Things that change the effective k
- With MMR,
kis what you keep andfetch_kis what you consider. - With post-filtering stores, a filter applied after the search can leave fewer than k results.
- With overlapping chunks, several of the k may be near-duplicates, so you see fewer distinct facts than k suggests.
A real-life example
An e-commerce product Q&A team uses k = 4 over reviews and spec chunks. For "How is the battery life compared with the previous model?", the answer needs both the spec line and two reviews that compare models. With k = 4 the spec line comes back but the comparison reviews rank 6th and 9th, so the answer ignores the comparison.
They measure on 150 labelled questions: recall@4 is 0.72, recall@10 is 0.86, and recall@30 is 0.95. They switch to retrieving 30, reranking, and sending the top 6. Recall of what the model actually sees rises close to the recall@30 level, and prompt size grows by only two chunks. Latency rises by the reranker's time, which they accept.
Follow-up questions to expect
- "How do you pick k for the first stage?" — Increase it until recall@k stops improving meaningfully on your labelled set, then check the reranker can score that many within your latency budget.
- "Why not send all 30 chunks to a long-context model?" — Cost and quality. It is 5 to 10 times more tokens, and irrelevant chunks can pull the answer in the wrong direction.
- "What k for multi-part questions?" — Larger, or better, split the question into sub-questions and retrieve for each.