Course Content
Advanced RAG
3 sections · 38 lessons
How do you decide how many retrieved documents to re-rank?
What you need to know
Two stages, two jobs
- First stage (BM25 + dense, fused with RRF): fast over millions of chunks; judged by recall@N.
- Reranker: slow per item, accurate; judged by precision at the top (for example, is the right chunk in the top 3?).
The method: find the knee
Record, for each question in your labelled set, the rank at which the correct chunk appears in first-stage results. Then compute recall@N for several N:
1# rank of the gold chunk in first-stage retrieval, one per eval question2# (None = not in the top 500 at all)3gold_ranks = [1, 3, 2, 14, 7, 41, 1, 88, 5, 23, 2, 160, 9, 1, None, 36, 4, 57, 12, 3]45for n in (10, 25, 50, 100, 200):6 hit = sum(r is not None and r <= n for r in gold_ranks)7 print(f"recall@{n:<3} = {hit / len(gold_ranks):.2f}")recall@10 = 0.55recall@25 = 0.70recall@50 = 0.80recall@100 = 0.90recall@200 = 0.95Each doubling buys less: 50 → 100 adds 10 points, 100 → 200 adds 5 and doubles rerank cost. Here N = 100 is a reasonable choice if latency allows, 50 if it doesn't. Real sets need hundreds of questions, but the shape is usually like this.
Latency and the cascade
A cross-encoder scores each (query, passage) pair separately, so cost is linear in N and in passage length. Batching on a GPU helps throughput but does not change that. If 100 candidates is too slow:
- Cheap reranker — a small cross-encoder or ColBERT-style late interaction over 100–200.
- Strong reranker — a large cross-encoder over the top 20–30.
- Optional LLM reranker — over the top 10, for hard or high-value queries.
LLM rerankers
LLMs can rerank too. Pointwise: score each passage separately ("rate relevance 0–3"). Listwise: show the model 10–20 passages and ask for an ordering (RankGPT popularised this; long lists are processed in sliding windows). LLM rerankers handle nuanced relevance and instructions ("prefer the current version") well, but cost more and add latency. Listwise ranking is also sensitive to the order in which passages are shown, so shuffle or run more than once for important cases. Several dedicated reranker models today are themselves built on small LLMs, which blurs the line.
Details people miss
- Cap by tokens, not just count. 100 chunks of 2,000 tokens cost far more than 100 of 200.
- Keep the score. A top reranker score below a floor means "nothing relevant found" — answer that instead of hallucinating.
- Tune the final top-k separately. How many reranked passages the generator gets (often 3–8) is a different question from how many are reranked.
A real-life example
A pharma company's regulatory search uses hybrid retrieval and a hosted cross-encoder. They currently rerank 25 candidates and send the top 5 to the model.
Scientists report that specific questions — "retest period extension requirements for a drug substance in Japan" — often miss. The team labels 300 questions and plots first-stage recall: 0.71 at N = 25, 0.86 at N = 75, 0.89 at N = 150. The misses were not reranker failures; the right passage was ranked 30th–70th by the first stage and never reached the reranker.
They raise N to 75. Rerank latency rises by a few hundred milliseconds, which fits the 3-second target. For the 10% of questions flagged as "comparison across regions", they add an LLM listwise reranker over the top 15 to prefer the current guideline versions. Final top-5 accuracy improves clearly, and the reranker's score floor now catches most questions whose answer is genuinely not in the corpus.
Follow-up questions to expect
- "Why not rerank everything?" — Cost and latency are linear in N, and recall gains flatten; past the knee you pay a lot for almost nothing.
- "Can the reranker make results worse?" — Yes, on domains it wasn't trained for, such as code or non-English text. Always compare reranked versus first-stage order on your own set; fine-tune or switch models if needed.
- "Cross-encoder or LLM reranker?" — Cross-encoder for most traffic: fast and strong. LLM reranker for a small, hard or high-value slice, where its better judgement is worth the cost.