Course Content
RAG Systems
12 sections · 66 lessons
What is reranking and why is it important in advanced RAG pipelines?
What you need to know
Bi-encoder (first stage)
- Embeds query and document separately
- Document vectors computed once, at ingest
- Compare with one dot product: millisecond search over millions
- Cannot see how query words relate to document words
Cross-encoder (reranker)
- Reads query and document together in one pass
- Nothing precomputed: one model call per pair
- 30 pairs per query is fine; millions is not
- Sees word-level interactions, so judges relevance better
In code
1from sentence_transformers import CrossEncoder23reranker = CrossEncoder("cross-encoder/ms-marco-MiniLM-L6-v2")4candidates = retriever.invoke(query) # e.g. 30 from hybrid search5scores = reranker.predict([(query, d.page_content) for d in candidates])6top = [d for _, d in sorted(zip(scores, candidates), key=lambda p: -p[0])][:4]Cross-encoder scores are raw numbers, not probabilities (in our runs this model gave about +6 for a clear match and about −11 for unrelated text), but they separate relevant from irrelevant more sharply than cosine scores, which makes a "nothing relevant" threshold easier to set.
What it did on a small test
On an HR handbook split into 45 small chunks, with 14 labelled questions, bge-small-en-v1.5 retrieved the top 10 and the cross-encoder above reordered them:
| Recall@1 | MRR | |
|---|---|---|
| Bi-encoder only | 10 / 14 | 0.80 |
| After reranking top 10 | 12 / 14 | 0.89 |
"Is there extra pay if I work long weeks?" had the overtime chunk at rank 7; after reranking it was at rank 2. This is a toy set, so treat it as an illustration, not a benchmark. One question failed in both, because the answer sentence had been split across two chunks; a reranker cannot fix chunking.
nDCG: when order and degree both matter
For "Will I be paid for unused leave when I quit?", grades: Leave encashment = 3, Carry forward = 1, Earned leave = 1, others 0.
Discounts for ranks 1..5: 1.000 0.631 0.500 0.431 0.387Before rerank: [1, 0, 3, 1, 0]DCG = 1(1.000) + 0 + 3(0.500) + 1(0.431) + 0 = 2.931Ideal: [3, 1, 1, 0, 0]IDCG = 3(1.000) + 1(0.631) + 1(0.500) = 4.131nDCG = 2.931 / 4.131 = 0.71 after rerank to ideal order: 1.00Options in 2026
- Open cross-encoders (the MiniLM MS MARCO models,
bge-rerankermodels) run on a CPU for small batches or a GPU at scale. - Hosted rerank APIs (Cohere, Voyage, Jina and others) need no infrastructure.
- LLM rerankers ask a language model to score or order candidates; they are strong but slower and more expensive.
- Late-interaction models (ColBERT family) sit between the two: token-level matching with precomputed document vectors, fast enough for first-stage search.
A real-life example
A legal-contract search tool retrieves 40 hybrid candidates per question. Recall@40 on 100 labelled questions is high, but the top 5 sent to the model often contain definitions and boilerplate that mention the right words, while the operative clause sits at rank 12.
They add a cross-encoder on a small GPU, rerank the 40 and send the top 5. The operative clause now reaches the model for most of those questions, and answer accuracy judged by associates improves. The reranker adds about 120 ms per question in their setup; they cap candidates at 40 and cache results by query and document IDs for repeated questions.
Follow-up questions to expect
- "Why not use the cross-encoder for everything?" — It must run once per query-document pair. For a million chunks that is a million model calls per question.
- "Can the reranker's score gate the answer?" — Yes. If the best reranked score is below a calibrated threshold, reply that the documents do not cover the question.
- "Does reranking help if recall@30 is low?" — No. It can only reorder what it is given; fix first-stage retrieval first.