RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

What does the k parameter control in vector retrieval?


Where the comparison review rankedspecrevfaqspecrevrevcomparerevfaqrev0123456789k = 4stops hereneeded, rank 7Retrieve 30, rerank, send 6: the model sees what recall@30 found.
A small k is not cheaper if the answer sits at rank 7 — retrieve wide, then rerank narrow.

What you need to know

k changes recall directly

If the right chunk is at rank 7, then recall@5 is 0 and recall@10 is 1 for that question. Across a labelled set, recall@k always rises as k rises, but with diminishing returns. A typical curve rises steeply to about k = 10, then flattens.

k changes cost directly

With 400-token chunks:

kContext tokensEffect
31,200Cheap, fast; misses split answers
52,000Common default for direct questions
104,000Better for comparisons and multi-part questions
3012,000Too much to send directly; fine as reranker input

At 10,000 questions a day, going from k = 5 to k = 30 adds 100 million input tokens a day. And more context is not always better for quality: irrelevant chunks distract the model, and facts in the middle of long prompts are used less reliably.

Two ks, not one

  1. Retrieve wide — hybrid search returns 30 candidates. This step is about recall: is the answer anywhere in the list?
  2. Rerank — a cross-encoder scores all 30 against the question.
  3. Send narrow — the top 3 to 5 go into the prompt. This step is about precision.

Measure recall@30 for the first step and recall@5 after reranking. If recall@30 is low, fix retrieval (chunking, hybrid, filters). If recall@30 is high but recall@5 is low, a reranker will help.

Adaptive k

You do not always need a fixed count. Keep chunks until the score drops sharply or below a calibrated threshold, with a minimum and maximum. A simple factual question may use 2 chunks; a comparison may use 8.

Things that change the effective k

  • With MMR, k is what you keep and fetch_k is what you consider.
  • With post-filtering stores, a filter applied after the search can leave fewer than k results.
  • With overlapping chunks, several of the k may be near-duplicates, so you see fewer distinct facts than k suggests.

A real-life example

An e-commerce product Q&A team uses k = 4 over reviews and spec chunks. For "How is the battery life compared with the previous model?", the answer needs both the spec line and two reviews that compare models. With k = 4 the spec line comes back but the comparison reviews rank 6th and 9th, so the answer ignores the comparison.

They measure on 150 labelled questions: recall@4 is 0.72, recall@10 is 0.86, and recall@30 is 0.95. They switch to retrieving 30, reranking, and sending the top 6. Recall of what the model actually sees rises close to the recall@30 level, and prompt size grows by only two chunks. Latency rises by the reranker's time, which they accept.

Follow-up questions to expect

  • "How do you pick k for the first stage?" — Increase it until recall@k stops improving meaningfully on your labelled set, then check the reranker can score that many within your latency budget.
  • "Why not send all 30 chunks to a long-context model?" — Cost and quality. It is 5 to 10 times more tokens, and irrelevant chunks can pull the answer in the wrong direction.
  • "What k for multi-part questions?" — Larger, or better, split the question into sub-questions and retrieve for each.