Course Content
Retrieval-Augmented Generation (RAG)
4 sections · 8 lessons
Context Re-ranking and Filtering
A compliance team built a RAG assistant over 40,000 internal policy documents. It worked. Mostly. Then an auditor asked "Who signs off on a wire transfer above the standard limit?" and it answered, confidently, "the Head of Treasury."
The correct answer was "two signatories, one of whom must be a Director" — and it was in the corpus, in Payments-Authority-Matrix-v7.pdf. The retriever had found it, at rank 9. The pipeline passed the top 5 chunks to the model. Rank 9 never reached the prompt.
Nobody had broken anything. The retriever did its job: the right document was in the candidate set. The generator did its job: it answered faithfully from the context it was handed. The failure lived entirely in the gap between "we found it" and "we showed it to the model" — and that gap is where re-ranking and filtering live.
This is the most under-built stage in most RAG systems, and usually the cheapest place to buy a large accuracy gain.
Why retrieval alone is not enough
A vector retriever answers one narrow question extremely fast: which stored vectors point roughly the same way as this query vector? That design lets it scan two million chunks in eight milliseconds. It also limits it, in three specific ways.
The score is a blunt instrument. One embedding compresses a 400-word chunk into 768 numbers. A chunk about approving wire transfers and one about reversing them land in nearly the same place: the topic dominates the vector and the verb is a rounding error.
The query and the document never meet. The document was embedded months ago, alone, knowing nothing of what would be asked of it. The query is embedded now, alone. They are compared only after both have been flattened. Nothing ever reads them together.
Top-k is a guess. You pick k=5 because it fits the context budget. Sometimes the answer is at rank 1 and ranks 2-5 are noise you are paying for. Sometimes it is at rank 9. The retriever cannot tell you which case you are in.
Retrieval optimises for recall at scale — get the answer in the pile somewhere, cheaply. Re-ranking optimises for precision at the top — get it to position one. These are different jobs and one component cannot do both well.
The architecture that works is two-stage: retrieve wide and shallow, then re-rank narrow and deep.
query | v[ retriever ] --> 50-200 candidates fast, approximate, high recall | v[ re-ranker ] --> 5-10 candidates slow, exact, high precision | v[ filters ] --> final context dedupe, threshold, budget | v[ generator ]Bi-encoders and cross-encoders
The intuition, before the maths
You have 500 CVs and one job description.
A bi-encoder reduces every CV to a five-word summary in advance — "senior backend, Go, fintech, London" — reduces the job description the same way, and matches summaries. Fast enough for all 500. It also confuses "led a team of eight" with "worked in a team of eight", because both compress to roughly "team, eight".
A cross-encoder reads one CV and the job description side by side. It catches that this candidate's fintech experience was compliance, not payments — exactly what the job needs. You cannot do that 500 times before lunch.
So use the fast one to get from 500 to 50, and the careful one to get from 50 to 5.
The formal difference
A bi-encoder encodes query and document independently:
Because E(d) does not depend on q, every document embedding is computed once, offline, and stored. At query time you compute E(q) once and search. Cost is one forward pass plus an index lookup, whatever the corpus size.
A cross-encoder encodes them jointly:
Query and document tokens enter the same transformer, so attention can compare "above the standard limit" against "exceeding the threshold in Schedule B" token by token. Nothing is precomputable, because the input contains the query. Cost is one forward pass per document.
| Bi-encoder | Cross-encoder | |
|---|---|---|
| Input | Query and doc encoded separately | Query and doc encoded together |
| Output | Two vectors, compared by cosine | A single relevance score |
| Precomputable | Yes — index built offline | No — depends on the query |
| Forward passes per query | 1 | 1 per candidate |
| Cost over 2M chunks | ~8 ms (HNSW index) | Impossible — would be days |
| Cost over 50 candidates | negligible | ~120 ms on a small GPU |
| Typical accuracy on hard pairs | Good | Substantially better |
| Role | First stage: recall | Second stage: precision |
The arithmetic explains the architecture by itself. A cross-encoder over 2,000,000 chunks at 2.4 ms per pair is 4,800 seconds — eighty minutes — for one query. Over 50 candidates it is 0.12 seconds. That factor of 40,000 is the whole reason for two stages.
Cross-encoders in practice
Basic use
1from sentence_transformers import CrossEncoder23# ~22M params, 512-token limit, fast enough for interactive use.4reranker = CrossEncoder("cross-encoder/ms-marco-MiniLM-L-6-v2")56query = "Who signs off on a wire transfer above the standard limit?"7candidates = [8 "Wire transfers are processed daily at 14:00 GMT by the operations desk.",9 "Payments above the standard limit require sign-off from two signatories, "10 "at least one of whom must hold Director status.",11 "The Head of Treasury owns the wire transfer policy document.",12]1314scores = reranker.predict([(query, c) for c in candidates])15for score, text in sorted(zip(scores, candidates), reverse=True):16 print(f"{score:7.3f} {text[:70]}") 6.345 Payments above the standard limit require sign-off from two signatorie -3.695 The Head of Treasury owns the wire transfer policy document. -5.791 Wire transfers are processed daily at 14:00 GMT by the operations deskNotice the third candidate — "wire transfer", "policy", a named authority, and exactly the chunk that produced the wrong answer above. A bi-encoder scores it highly: with all-MiniLM-L6-v2 it gets a cosine of 0.53 against 0.51 for the correct chunk, so it ranks first. The cross-encoder, reading query and chunk together, sees that owning a policy document is not signing off on a transfer, and pushes it down.
Those numbers are logits, not probabilities
Most cross-encoders trained on MS MARCO output an unbounded logit, not a 0-to-1 score. To threshold on it, push it through a sigmoid:
For the scores above: σ(6.345)=0.9982, σ(−3.695)=0.0243, σ(−5.791)=0.0030. Comparable and thresholdable — but only within this model. A logit of 6.3 from MiniLM-L-6 says nothing about 6.3 from bge-reranker-large. Retune every threshold on every model change.
Wiring it into a retriever
1def retrieve_and_rerank(query, vectorstore, reranker,2 fetch_k=50, top_k=5):3 # Stage 1: wide and cheap. Ask for far more than you need.4 candidates = vectorstore.similarity_search(query, k=fetch_k)56 # Stage 2: narrow and careful.7 pairs = [(query, d.page_content) for d in candidates]8 scores = reranker.predict(pairs, batch_size=32)910 ranked = sorted(zip(scores, candidates), key=lambda x: x[0],11 reverse=True)12 return [(float(s), d) for s, d in ranked[:top_k]]fetch_k is the parameter that matters most, and setting it too low is the commonest mistake. A re-ranker only reorders what it is given. If the answer is not among the 50 candidates, no re-ranker will surface it. Recall at fetch_k is a hard ceiling on final accuracy, so measure it before tuning anything else. Illustrative numbers from one corpus:
| fetch_k | Recall@fetch_k | Recall@5 after re-rank | Re-rank latency |
|---|---|---|---|
| 10 | 0.68 | 0.64 | ~25 ms |
| 20 | 0.79 | 0.71 | ~50 ms |
| 50 | 0.91 | 0.83 | ~120 ms |
| 100 | 0.94 | 0.86 | ~240 ms |
| 200 | 0.96 | 0.87 | ~480 ms |
Read the last two rows: doubling from 100 to 200 costs 240 ms and buys one point. The knee sits around 50 for most corpora — find your own, do not inherit someone else's.
Multi-stage re-ranking
One pass is not the limit. Cascade cheap-to-expensive, so the expensive model only ever sees a handful of documents.
| Stage | Method | In → Out | Typical cost |
|---|---|---|---|
| 0 | Hybrid BM25 + dense | 2M → 200 | ~15 ms |
| 1 | Small cross-encoder (MiniLM-L-2) | 200 → 50 | ~90 ms |
| 2 | Strong cross-encoder (bge-reranker-large) | 50 → 8 | ~200 ms |
| 3 | LLM listwise re-rank | 8 → 5 | ~700 ms, ~0.01 in tokens |
Stage 3 shows the model several candidates at once and asks it to order them. Because it sees them together it can make comparative judgements a pairwise scorer cannot — "chunk C supersedes chunk A, it is the v7 revision". It is also the most expensive thing in the pipeline. Add it only once you have measured that stages 0-2 have plateaued. If you would rather not serve re-ranking models yourself, hosted re-rank APIs from providers such as Cohere and Voyage AI fill the same slot as stage 2.
Filtering: removing what should not be there
Re-ranking orders; filtering removes. Four filters earn their place.
Length filtering
Chunkers produce garbage at the edges: a chunk containing only "Section 4.2", a page number, a caption orphaned from its table. Short text has a sharp embedding and scores well while teaching the generator nothing. Drop anything under roughly 20 tokens of prose.
The opposite is also real: a 3,000-token chunk holding the answer in one sentence carries 2,900 tokens of distraction, and models degrade measurably when the fact sits buried mid-block.
Score threshold filtering
If the best re-ranked score is low, five weak chunks are worse than two decent ones — each extra weak chunk is another chance for the generator to anchor on something wrong.
1def threshold_filter(scored_docs, min_prob=0.30, min_docs=1):2 import math3 kept = [(s, d) for s, d in scored_docs4 if 1 / (1 + math.exp(-s)) >= min_prob]5 # Never return nothing: keep the single best so the generator6 # can at least say "the closest I found does not answer this".7 return kept if kept else scored_docs[:min_docs]The min_docs fallback matters: a threshold that can return an empty list produces a prompt with no context, and a model given no context answers from its parameters — precisely the hallucination you built RAG to prevent.
Redundancy filtering with MMR
Corpora are full of near-duplicates: the same policy paragraph in v6 and v7, the same FAQ answer on three pages. Five chunks that all say the same thing give the generator one fact and consume five slots. Maximal Marginal Relevance scores each candidate on relevance minus similarity to what is already picked:
where S is the set already selected. Work it through with λ=0.7 and four candidates:
Doc sim to query sim to A sim to C A 0.82 — 0.42 B 0.80 0.95 0.40 C 0.78 0.42 — D 0.61 0.30 0.35 Step 1: nothing selected, so take the highest relevance — A at 0.82.
Step 2: penalise similarity to A. B scores 0.7(0.80)−0.3(0.95)=0.275; C scores 0.7(0.78)−0.3(0.42)=0.420; D scores 0.7(0.61)−0.3(0.30)=0.337. Pick C, despite B having the higher raw similarity.
Step 3: the penalty is now the max against A and C. B stays at 0.560−0.3(0.95)=0.275; D gives 0.427−0.3(0.35)=0.322. Pick D.
Final order A, C, D, B against pure similarity's A, B, C, D. B, a 0.95 near-duplicate of A, is demoted to last, so three distinct facts reach the generator instead of two facts and an echo. Tune λ deliberately: 0.8 for factual lookup, 0.5 when the question is "summarise everything we know about X".
Metadata filtering — and where to apply it
Restricting by source, date, department or access level removes documents that could never have been correct. The critical detail is where it runs — filtering after retrieval is a bug that looks like a feature:
Python1# WRONG: 50 in, maybe 3 survive, and treasury docs at rank 51-3002# were never considered at all.3docs = store.similarity_search(query, k=50)4docs = [d for d in docs if d.metadata["dept"] == "treasury"]56# RIGHT: push the predicate into the index. 50 eligible candidates.7docs = store.similarity_search(8 query, k=50, filter={"dept": "treasury", "status": "current"}9)1011# Filter syntax is store-specific. Chroma, for example, allows one12# condition per dict, so two conditions must be combined with $and:13docs = chroma_store.similarity_search(14 query, k=50,15 filter={"$and": [{"dept": "treasury"}, {"status": "current"}]},16)pgvector, Qdrant, Weaviate and Pinecone all support pre-filtered search. Post-filtering silently shrinks the candidate set and destroys the recall you paid for.
Order the filters correctly
The sequence is not arbitrary: cheaper and safer before expensive and riskier.
Order Step Why here 1 Metadata pre-filter (in the index) Free, and removes ineligible docs before they consume candidate slots 2 Retrieve fetch_k candidates Wide net; recall ceiling is set here 3 Length filter Cheap; strips fragments before you pay to score them 4 Cross-encoder re-rank The expensive, accurate step — run it on a clean candidate set 5 Score threshold Now the scores are trustworthy enough to threshold on 6 MMR / redundancy Diversify among documents already known to be relevant 7 Context budget packing Last: fit what survived into the token allowance Never threshold on first-stage retriever scores. Those scores are the crude signal you brought the re-ranker in to correct — filtering on them throws away exactly the documents the re-ranker was going to promote.
Confidence-aware retrieval
With calibrated scores you can ask what most pipelines never ask: should we answer this at all? One score is a weak signal; three combined are much stronger.
- Top score — how good is the best evidence?
- Margin — how far above the fifth-best? A large gap means one document clearly answers this; a flat distribution means nothing does and the ranking is noise.
- Support count — how many chunks clear the bar? One is fragile, three agreeing is solid.
1import math23def sigmoid(x):4 return 1 / (1 + math.exp(-x))56def confidence(scores, bar=0.30):7 """scores: re-ranker logits, already sorted descending."""8 probs = [sigmoid(s) for s in scores]9 top = probs[0]10 margin = probs[0] - probs[min(4, len(probs) - 1)]11 support = sum(1 for p in probs if p >= bar) / max(len(probs), 1)12 return 0.5 * top + 0.3 * margin + 0.2 * supportTwo contrasting cases, same model:
| Query A (answerable) | Query B (not in corpus) | |
|---|---|---|
| Top-5 logits | 8.2, 3.1, −0.4, −2.0, −4.1 | 0.6, 0.4, 0.2, 0.1, −0.1 |
| Top-5 probabilities | 0.9997, 0.9569, 0.4013, 0.1192, 0.0163 | 0.6457, 0.5987, 0.5498, 0.5250, 0.4750 |
| Top | 0.9997 | 0.6457 |
| Margin (1st − 5th) | 0.9834 | 0.1707 |
| Support (≥ 0.30) | 3/5 = 0.60 | 5/5 = 1.00 |
| Confidence | 0.915 | 0.574 |
Query B is the dangerous one, and it fools the naive check: all five chunks clear the 0.30 bar, so support is perfect — better than Query A's. The margin gives it away, 0.17 against 0.98. Five equally, mildly related documents is what a corpus with nothing specific to say looks like in numbers.
Then act on it.
| Confidence | Behaviour |
|---|---|
| > 0.75 | Answer normally from context |
| 0.45 – 0.75 | Answer, but hedge and cite explicitly; widen fetch_k and retry once first |
| < 0.45 | Do not answer from context. Say the corpus does not cover this, offer the closest documents found, or hand off |
A RAG system that can say "I could not find this in your documents" is more valuable than one that is right slightly more often but never admits doubt. Users forgive a gap. They do not forgive a confident, well-formatted, wrong answer with a citation attached.
Measuring whether re-ranking helped
Precision and recall ignore order, which makes them useless here: re-ranking changes nothing but order. You need rank-sensitive metrics.
MRR — Mean Reciprocal Rank
Find the rank of the first relevant document, take its reciprocal, average over queries. Rank 1 scores 1.0, rank 2 scores 0.5, rank 4 scores 0.25, not found scores 0.
| Query | Rank before | RR before | Rank after | RR after |
|---|---|---|---|---|
| 1 | 1 | 1.0000 | 1 | 1.0000 |
| 2 | 3 | 0.3333 | 1 | 1.0000 |
| 3 | 2 | 0.5000 | 2 | 0.5000 |
| 4 | not found | 0.0000 | 4 | 0.2500 |
| 5 | 5 | 0.2000 | 1 | 1.0000 |
| Sum | 2.0333 | 3.7500 |
MRR before: 2.0333/5=0.4067. After: 3.75/5=0.7500 — an 84% relative gain purely from reordering documents the retriever had already found. Query 4 is the opening scenario, fixed. MRR is right when there is one correct answer; it is wrong when several documents are relevant, because it ignores everything after the first hit.
NDCG@K — Normalised Discounted Cumulative Gain
NDCG handles graded relevance (3 = perfect, 2 = useful, 1 = marginal, 0 = irrelevant) and rewards putting the best material highest. Gain at rank i is discounted by log2(i+1):
Say the retriever returns grades [0, 2, 0, 3, 1] at ranks 1-5:
| Rank i | reli | log2(i+1) | Contribution |
|---|---|---|---|
| 1 | 0 | 1.0000 | 0.0000 |
| 2 | 2 | 1.5850 | 1.2619 |
| 3 | 0 | 2.0000 | 0.0000 |
| 4 | 3 | 2.3219 | 1.2921 |
| 5 | 1 | 2.5850 | 0.3868 |
| DCG@5 | 2.9408 |
The ideal ordering of those same grades is [3, 2, 1, 0, 0]:
IDCG@5=1.00003+1.58502+2.00001=3.0000+1.2619+0.5000=4.7619
So NDCG@5=2.9408/4.7619=0.618.
Now re-rank to [3, 2, 0, 1, 0]: DCG@5=3.0000+1.2619+0+2.32191+0=4.6926, giving NDCG@5=4.6926/4.7619=0.985.
0.618 to 0.985 — same five documents, different order.
One trap: two gain formulas are in common use. The linear one above uses reli; the exponential one uses 2reli−1, weighting a grade-3 document seven times a grade-1 rather than three times. On these same lists it gives 0.564 before and 0.993 after. Both are called "NDCG@5", so always state which you used and never compare across conventions.
The latency and accuracy trade-off
Every re-ranking decision spends milliseconds to buy precision. Whether that trade is good depends entirely on what the system is for.
| Configuration | Added latency | NDCG@5 (illustrative) | Fits |
|---|---|---|---|
| No re-ranking | 0 ms | 0.61 | Autocomplete, live suggestions |
| MiniLM-L-6, fetch_k=20 | +50 ms | 0.74 | Interactive chat — the default choice |
| MiniLM-L-6, fetch_k=50 | +120 ms | 0.83 | Interactive chat where accuracy matters |
| bge-reranker-large, fetch_k=50 | +320 ms | 0.88 | Support, research, internal tools |
| Cascade + LLM listwise | +1,100 ms | 0.92 | Legal, medical, compliance; batch jobs |
Budget against the whole request, not the re-ranker alone. If generation already takes 900 ms, a 120 ms re-ranker adds 13% to perceived latency and lifts NDCG by twenty-two points — not a close call. Serving typeahead in 40 ms, the same re-ranker is a 300% increase and obviously wrong.
Three levers to claw latency back, cheapest first: cut fetch_k (linear saving, flat curve above the knee); drop to a smaller re-ranker (MiniLM-L-2 is about half the cost of L-6); cache scores, which are deterministic per query-document pair.
Where this goes wrong
Six failure modes account for nearly everything that breaks here.
| Symptom | Cause | Fix |
|---|---|---|
| Re-ranking barely moves accuracy | fetch_k too small — the answer was never in the candidate set | Measure recall@fetch_k first; raise fetch_k until it plateaus |
| Long chunks score inexplicably low | Cross-encoders truncate at 512 tokens; the relevant sentence was cut off | Chunk under ~400 tokens, or score a summary/window of the chunk |
| Thresholds stop working after a model swap | Logit scales are model-specific and uncalibrated | Re-tune thresholds on a held-out set for every model change |
| Retrieval returns 2 docs where 50 were expected | Metadata filter applied after retrieval instead of inside the index | Push the predicate into the vector store's filtered search |
| Five chunks all say the same thing | No redundancy filter; corpus has versioned near-duplicates | MMR with λ≈0.7, and deduplicate at ingestion |
| Confident wrong answers on out-of-scope questions | No confidence gate; a flat score profile was treated as a good result | Score top + margin + support; refuse below threshold |
The first is subtlest and worth naming plainly: re-ranking cannot create recall. If the top 50 contains the answer 68% of the time, 68% is your absolute accuracy ceiling however good the re-ranker is. Teams routinely spend a fortnight benchmarking re-rankers when the real problem is a chunker that split the answer across two chunks so neither holds it whole.
What this means when you build one
Start by measuring the ceiling. Take 50 real questions with known answers, retrieve at k=50, and record how often the answer appears anywhere in that list. That one number tells you which problem you have: below about 0.85 it is chunking, embeddings or missing keyword search, and a re-ranker will not help; above 0.85 it is ordering, and a re-ranker is the highest-return change available.
When it is an ordering problem, add the smallest thing that works: ms-marco-MiniLM-L-6-v2, fetch_k=50, top_k=5. Three lines and about 120 ms. Measure NDCG@5 and MRR before and after on the same 50 questions; you want a jump of ten points or more. If you do not get it, the ceiling was your problem after all.
Then add filters in order: metadata pre-filtering inside the index, a score threshold with a floor guaranteeing one survivor, and MMR at λ=0.7 if your corpus has versioned documents — it almost certainly does.
Last, add the confidence gate and make it visible: log the score on every request alongside the query. Within a week you will have a ranked list of the questions your corpus genuinely cannot answer — the most actionable document a RAG team can own, because it tells you what to write, what to ingest, and what to stop pretending you cover. The compliance team above found their wire-transfer failure that way, not from a bug report but from a cluster of low-margin queries all circling the same missing authority matrix.