RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

What is reranking and why is it important in advanced RAG pipelines?


Why the cross-encoder goes secondBi-encoder, first stage• Query and chunk embedded separately• Chunk vectors precomputed at ingest• Millions searched in milliseconds• Overtime chunk ranked 7thCross-encoder, reranker• Query and chunk read together• One model call per pair• Only 10 to 50 candidates• Overtime chunk moved to 2nd
On 14 toy questions MRR rose from 0.80 to 0.89 — but the one answer split across two chunks stayed lost.

What you need to know

Bi-encoder (first stage)

  • Embeds query and document separately
  • Document vectors computed once, at ingest
  • Compare with one dot product: millisecond search over millions
  • Cannot see how query words relate to document words

Cross-encoder (reranker)

  • Reads query and document together in one pass
  • Nothing precomputed: one model call per pair
  • 30 pairs per query is fine; millions is not
  • Sees word-level interactions, so judges relevance better

In code

Python
from sentence_transformers import CrossEncoderreranker = CrossEncoder("cross-encoder/ms-marco-MiniLM-L6-v2")candidates = retriever.invoke(query)                  # e.g. 30 from hybrid searchscores = reranker.predict([(query, d.page_content) for d in candidates])top = [d for _, d in sorted(zip(scores, candidates), key=lambda p: -p[0])][:4]

Cross-encoder scores are raw numbers, not probabilities (in our runs this model gave about +6 for a clear match and about −11 for unrelated text), but they separate relevant from irrelevant more sharply than cosine scores, which makes a "nothing relevant" threshold easier to set.

What it did on a small test

On an HR handbook split into 45 small chunks, with 14 labelled questions, bge-small-en-v1.5 retrieved the top 10 and the cross-encoder above reordered them:

Recall@1MRR
Bi-encoder only10 / 140.80
After reranking top 1012 / 140.89

"Is there extra pay if I work long weeks?" had the overtime chunk at rank 7; after reranking it was at rank 2. This is a toy set, so treat it as an illustration, not a benchmark. One question failed in both, because the answer sentence had been split across two chunks; a reranker cannot fix chunking.

nDCG: when order and degree both matter

For "Will I be paid for unused leave when I quit?", grades: Leave encashment = 3, Carry forward = 1, Earned leave = 1, others 0.

Text
Discounts for ranks 1..5:  1.000  0.631  0.500  0.431  0.387Before rerank: [1, 0, 3, 1, 0]DCG  = 1(1.000) + 0 + 3(0.500) + 1(0.431) + 0           = 2.931Ideal: [3, 1, 1, 0, 0]IDCG = 3(1.000) + 1(0.631) + 1(0.500)                     = 4.131nDCG = 2.931 / 4.131 = 0.71        after rerank to ideal order: 1.00

Options in 2026

  • Open cross-encoders (the MiniLM MS MARCO models, bge-reranker models) run on a CPU for small batches or a GPU at scale.
  • Hosted rerank APIs (Cohere, Voyage, Jina and others) need no infrastructure.
  • LLM rerankers ask a language model to score or order candidates; they are strong but slower and more expensive.
  • Late-interaction models (ColBERT family) sit between the two: token-level matching with precomputed document vectors, fast enough for first-stage search.

A real-life example

A legal-contract search tool retrieves 40 hybrid candidates per question. Recall@40 on 100 labelled questions is high, but the top 5 sent to the model often contain definitions and boilerplate that mention the right words, while the operative clause sits at rank 12.

They add a cross-encoder on a small GPU, rerank the 40 and send the top 5. The operative clause now reaches the model for most of those questions, and answer accuracy judged by associates improves. The reranker adds about 120 ms per question in their setup; they cap candidates at 40 and cache results by query and document IDs for repeated questions.

Follow-up questions to expect

  • "Why not use the cross-encoder for everything?" — It must run once per query-document pair. For a million chunks that is a million model calls per question.
  • "Can the reranker's score gate the answer?" — Yes. If the best reranked score is below a calibrated threshold, reply that the documents do not cover the question.
  • "Does reranking help if recall@30 is low?" — No. It can only reorder what it is given; fix first-stage retrieval first.