Retrieval-Augmented Generation (RAG)

Course Content

Retrieval-Augmented Generation (RAG)

4 sections · 8 lessons

Context Re-ranking and Filtering


A compliance team built a RAG assistant over 40,000 internal policy documents. It worked. Mostly. Then an auditor asked "Who signs off on a wire transfer above the standard limit?" and it answered, confidently, "the Head of Treasury."

The correct answer was "two signatories, one of whom must be a Director" — and it was in the corpus, in Payments-Authority-Matrix-v7.pdf. The retriever had found it, at rank 9. The pipeline passed the top 5 chunks to the model. Rank 9 never reached the prompt.

Nobody had broken anything. The retriever did its job: the right document was in the candidate set. The generator did its job: it answered faithfully from the context it was handed. The failure lived entirely in the gap between "we found it" and "we showed it to the model" — and that gap is where re-ranking and filtering live.

This is the most under-built stage in most RAG systems, and usually the cheapest place to buy a large accuracy gain.

Bi-encoder top 8, before the cross-encoder sees it0.810.790.780.770.760.750.740.7101234567sent tothe modelthe signing ruleEight scores inside 0.10 of each other: the bi-encoder has ranked them, but it has not separated them.
A cross-encoder reads query and chunk together and re-scores these eight in about 40 ms, which is the only cheap way to move rank 7 to rank 1 — and it can never recover a chunk retrieval left out.

Why retrieval alone is not enough

A vector retriever answers one narrow question extremely fast: which stored vectors point roughly the same way as this query vector? That design lets it scan two million chunks in eight milliseconds. It also limits it, in three specific ways.

The score is a blunt instrument. One embedding compresses a 400-word chunk into 768 numbers. A chunk about approving wire transfers and one about reversing them land in nearly the same place: the topic dominates the vector and the verb is a rounding error.

The query and the document never meet. The document was embedded months ago, alone, knowing nothing of what would be asked of it. The query is embedded now, alone. They are compared only after both have been flattened. Nothing ever reads them together.

Top-k is a guess. You pick k=5 because it fits the context budget. Sometimes the answer is at rank 1 and ranks 2-5 are noise you are paying for. Sometimes it is at rank 9. The retriever cannot tell you which case you are in.

Retrieval optimises for recall at scale — get the answer in the pile somewhere, cheaply. Re-ranking optimises for precision at the top — get it to position one. These are different jobs and one component cannot do both well.

The architecture that works is two-stage: retrieve wide and shallow, then re-rank narrow and deep.

Text
query  |  v[ retriever ]  --> 50-200 candidates   fast, approximate, high recall  |  v[ re-ranker ]  --> 5-10 candidates     slow, exact, high precision  |  v[ filters ]    --> final context       dedupe, threshold, budget  |  v[ generator ]

Bi-encoders and cross-encoders

The intuition, before the maths

You have 500 CVs and one job description.

A bi-encoder reduces every CV to a five-word summary in advance — "senior backend, Go, fintech, London" — reduces the job description the same way, and matches summaries. Fast enough for all 500. It also confuses "led a team of eight" with "worked in a team of eight", because both compress to roughly "team, eight".

A cross-encoder reads one CV and the job description side by side. It catches that this candidate's fintech experience was compliance, not payments — exactly what the job needs. You cannot do that 500 times before lunch.

So use the fast one to get from 500 to 50, and the careful one to get from 50 to 5.

The formal difference

A bi-encoder encodes query and document independently:

score(q,d)=cos⁡(E(q), E(d))\text{score}(q,d) = \cos\big(E(q),\, E(d)\big)

Because E(d)E(d) does not depend on qq, every document embedding is computed once, offline, and stored. At query time you compute E(q)E(q) once and search. Cost is one forward pass plus an index lookup, whatever the corpus size.

A cross-encoder encodes them jointly:

score(q,d)=f([CLS]  q  [SEP]  d)\text{score}(q,d) = f\big([\texttt{CLS}]\;q\;[\texttt{SEP}]\;d\big)

Query and document tokens enter the same transformer, so attention can compare "above the standard limit" against "exceeding the threshold in Schedule B" token by token. Nothing is precomputable, because the input contains the query. Cost is one forward pass per document.

Bi-encoderCross-encoder
InputQuery and doc encoded separatelyQuery and doc encoded together
OutputTwo vectors, compared by cosineA single relevance score
PrecomputableYes — index built offlineNo — depends on the query
Forward passes per query11 per candidate
Cost over 2M chunks~8 ms (HNSW index)Impossible — would be days
Cost over 50 candidatesnegligible~120 ms on a small GPU
Typical accuracy on hard pairsGoodSubstantially better
RoleFirst stage: recallSecond stage: precision

The arithmetic explains the architecture by itself. A cross-encoder over 2,000,000 chunks at 2.4 ms per pair is 4,800 seconds — eighty minutes — for one query. Over 50 candidates it is 0.12 seconds. That factor of 40,000 is the whole reason for two stages.

Cross-encoders in practice

Basic use

Python
from sentence_transformers import CrossEncoder# ~22M params, 512-token limit, fast enough for interactive use.reranker = CrossEncoder("cross-encoder/ms-marco-MiniLM-L-6-v2")query = "Who signs off on a wire transfer above the standard limit?"candidates = [    "Wire transfers are processed daily at 14:00 GMT by the operations desk.",    "Payments above the standard limit require sign-off from two signatories, "    "at least one of whom must hold Director status.",    "The Head of Treasury owns the wire transfer policy document.",]scores = reranker.predict([(query, c) for c in candidates])for score, text in sorted(zip(scores, candidates), reverse=True):    print(f"{score:7.3f}  {text[:70]}")
Text
  6.345  Payments above the standard limit require sign-off from two signatorie -3.695  The Head of Treasury owns the wire transfer policy document. -5.791  Wire transfers are processed daily at 14:00 GMT by the operations desk

Notice the third candidate — "wire transfer", "policy", a named authority, and exactly the chunk that produced the wrong answer above. A bi-encoder scores it highly: with all-MiniLM-L6-v2 it gets a cosine of 0.53 against 0.51 for the correct chunk, so it ranks first. The cross-encoder, reading query and chunk together, sees that owning a policy document is not signing off on a transfer, and pushes it down.

Those numbers are logits, not probabilities

Most cross-encoders trained on MS MARCO output an unbounded logit, not a 0-to-1 score. To threshold on it, push it through a sigmoid:

σ(x)=11+e−x\sigma(x) = \frac{1}{1 + e^{-x}}

For the scores above: σ(6.345)=0.9982\sigma(6.345) = 0.9982, σ(−3.695)=0.0243\sigma(-3.695) = 0.0243, σ(−5.791)=0.0030\sigma(-5.791) = 0.0030. Comparable and thresholdable — but only within this model. A logit of 6.3 from MiniLM-L-6 says nothing about 6.3 from bge-reranker-large. Retune every threshold on every model change.

Wiring it into a retriever

Python
def retrieve_and_rerank(query, vectorstore, reranker,                        fetch_k=50, top_k=5):    # Stage 1: wide and cheap. Ask for far more than you need.    candidates = vectorstore.similarity_search(query, k=fetch_k)    # Stage 2: narrow and careful.    pairs = [(query, d.page_content) for d in candidates]    scores = reranker.predict(pairs, batch_size=32)    ranked = sorted(zip(scores, candidates), key=lambda x: x[0],                    reverse=True)    return [(float(s), d) for s, d in ranked[:top_k]]

fetch_k is the parameter that matters most, and setting it too low is the commonest mistake. A re-ranker only reorders what it is given. If the answer is not among the 50 candidates, no re-ranker will surface it. Recall at fetch_k is a hard ceiling on final accuracy, so measure it before tuning anything else. Illustrative numbers from one corpus:

fetch_kRecall@fetch_kRecall@5 after re-rankRe-rank latency
100.680.64~25 ms
200.790.71~50 ms
500.910.83~120 ms
1000.940.86~240 ms
2000.960.87~480 ms

Read the last two rows: doubling from 100 to 200 costs 240 ms and buys one point. The knee sits around 50 for most corpora — find your own, do not inherit someone else's.

Multi-stage re-ranking

One pass is not the limit. Cascade cheap-to-expensive, so the expensive model only ever sees a handful of documents.

StageMethodIn → OutTypical cost
0Hybrid BM25 + dense2M → 200~15 ms
1Small cross-encoder (MiniLM-L-2)200 → 50~90 ms
2Strong cross-encoder (bge-reranker-large)50 → 8~200 ms
3LLM listwise re-rank8 → 5~700 ms, ~0.01 in tokens

Stage 3 shows the model several candidates at once and asks it to order them. Because it sees them together it can make comparative judgements a pairwise scorer cannot — "chunk C supersedes chunk A, it is the v7 revision". It is also the most expensive thing in the pipeline. Add it only once you have measured that stages 0-2 have plateaued. If you would rather not serve re-ranking models yourself, hosted re-rank APIs from providers such as Cohere and Voyage AI fill the same slot as stage 2.

Filtering: removing what should not be there

Re-ranking orders; filtering removes. Four filters earn their place.

Length filtering

Chunkers produce garbage at the edges: a chunk containing only "Section 4.2", a page number, a caption orphaned from its table. Short text has a sharp embedding and scores well while teaching the generator nothing. Drop anything under roughly 20 tokens of prose.

The opposite is also real: a 3,000-token chunk holding the answer in one sentence carries 2,900 tokens of distraction, and models degrade measurably when the fact sits buried mid-block.

Score threshold filtering

If the best re-ranked score is low, five weak chunks are worse than two decent ones — each extra weak chunk is another chance for the generator to anchor on something wrong.

Python
def threshold_filter(scored_docs, min_prob=0.30, min_docs=1):    import math    kept = [(s, d) for s, d in scored_docs            if 1 / (1 + math.exp(-s)) >= min_prob]    # Never return nothing: keep the single best so the generator    # can at least say "the closest I found does not answer this".    return kept if kept else scored_docs[:min_docs]

The min_docs fallback matters: a threshold that can return an empty list produces a prompt with no context, and a model given no context answers from its parameters — precisely the hallucination you built RAG to prevent.

Redundancy filtering with MMR

Corpora are full of near-duplicates: the same policy paragraph in v6 and v7, the same FAQ answer on three pages. Five chunks that all say the same thing give the generator one fact and consume five slots. Maximal Marginal Relevance scores each candidate on relevance minus similarity to what is already picked:

MMR(d)=λ⋅sim(d,q)−(1−λ)⋅max⁡d′∈Ssim(d,d′)\text{MMR}(d) = \lambda \cdot \text{sim}(d, q) - (1-\lambda) \cdot \max_{d' \in S} \text{sim}(d, d')

where SS is the set already selected. Work it through with λ=0.7\lambda = 0.7 and four candidates:

Docsim to querysim to Asim to C
A0.82—0.42
B0.800.950.40
C0.780.42—
D0.610.300.35

Step 1: nothing selected, so take the highest relevance — A at 0.82.

Step 2: penalise similarity to A. B scores 0.7(0.80)−0.3(0.95)=0.2750.7(0.80) - 0.3(0.95) = 0.275; C scores 0.7(0.78)−0.3(0.42)=0.4200.7(0.78) - 0.3(0.42) = 0.420; D scores 0.7(0.61)−0.3(0.30)=0.3370.7(0.61) - 0.3(0.30) = 0.337. Pick C, despite B having the higher raw similarity.

Step 3: the penalty is now the max against A and C. B stays at 0.560−0.3(0.95)=0.2750.560 - 0.3(0.95) = 0.275; D gives 0.427−0.3(0.35)=0.3220.427 - 0.3(0.35) = 0.322. Pick D.

Final order A, C, D, B against pure similarity's A, B, C, D. B, a 0.95 near-duplicate of A, is demoted to last, so three distinct facts reach the generator instead of two facts and an echo. Tune λ\lambda deliberately: 0.8 for factual lookup, 0.5 when the question is "summarise everything we know about X".

Metadata filtering — and where to apply it

Restricting by source, date, department or access level removes documents that could never have been correct. The critical detail is where it runs — filtering after retrieval is a bug that looks like a feature:

Python
# WRONG: 50 in, maybe 3 survive, and treasury docs at rank 51-300# were never considered at all.docs = store.similarity_search(query, k=50)docs = [d for d in docs if d.metadata["dept"] == "treasury"]# RIGHT: push the predicate into the index. 50 eligible candidates.docs = store.similarity_search(    query, k=50, filter={"dept": "treasury", "status": "current"})# Filter syntax is store-specific. Chroma, for example, allows one# condition per dict, so two conditions must be combined with $and:docs = chroma_store.similarity_search(    query, k=50,    filter={"$and": [{"dept": "treasury"}, {"status": "current"}]},)

pgvector, Qdrant, Weaviate and Pinecone all support pre-filtered search. Post-filtering silently shrinks the candidate set and destroys the recall you paid for.

Order the filters correctly

The sequence is not arbitrary: cheaper and safer before expensive and riskier.

OrderStepWhy here
1Metadata pre-filter (in the index)Free, and removes ineligible docs before they consume candidate slots
2Retrieve fetch_k candidatesWide net; recall ceiling is set here
3Length filterCheap; strips fragments before you pay to score them
4Cross-encoder re-rankThe expensive, accurate step — run it on a clean candidate set
5Score thresholdNow the scores are trustworthy enough to threshold on
6MMR / redundancyDiversify among documents already known to be relevant
7Context budget packingLast: fit what survived into the token allowance

Never threshold on first-stage retriever scores. Those scores are the crude signal you brought the re-ranker in to correct — filtering on them throws away exactly the documents the re-ranker was going to promote.

Confidence-aware retrieval

With calibrated scores you can ask what most pipelines never ask: should we answer this at all? One score is a weak signal; three combined are much stronger.

  • Top score — how good is the best evidence?
  • Margin — how far above the fifth-best? A large gap means one document clearly answers this; a flat distribution means nothing does and the ranking is noise.
  • Support count — how many chunks clear the bar? One is fragile, three agreeing is solid.
Python
import mathdef sigmoid(x):    return 1 / (1 + math.exp(-x))def confidence(scores, bar=0.30):    """scores: re-ranker logits, already sorted descending."""    probs = [sigmoid(s) for s in scores]    top     = probs[0]    margin  = probs[0] - probs[min(4, len(probs) - 1)]    support = sum(1 for p in probs if p >= bar) / max(len(probs), 1)    return 0.5 * top + 0.3 * margin + 0.2 * support

Two contrasting cases, same model:

Query A (answerable)Query B (not in corpus)
Top-5 logits8.2, 3.1, −0.4, −2.0, −4.10.6, 0.4, 0.2, 0.1, −0.1
Top-5 probabilities0.9997, 0.9569, 0.4013, 0.1192, 0.01630.6457, 0.5987, 0.5498, 0.5250, 0.4750
Top0.99970.6457
Margin (1st − 5th)0.98340.1707
Support (≥ 0.30)3/5 = 0.605/5 = 1.00
Confidence0.9150.574

Query B is the dangerous one, and it fools the naive check: all five chunks clear the 0.30 bar, so support is perfect — better than Query A's. The margin gives it away, 0.17 against 0.98. Five equally, mildly related documents is what a corpus with nothing specific to say looks like in numbers.

Then act on it.

ConfidenceBehaviour
> 0.75Answer normally from context
0.45 – 0.75Answer, but hedge and cite explicitly; widen fetch_k and retry once first
< 0.45Do not answer from context. Say the corpus does not cover this, offer the closest documents found, or hand off

A RAG system that can say "I could not find this in your documents" is more valuable than one that is right slightly more often but never admits doubt. Users forgive a gap. They do not forgive a confident, well-formatted, wrong answer with a citation attached.

Measuring whether re-ranking helped

Precision and recall ignore order, which makes them useless here: re-ranking changes nothing but order. You need rank-sensitive metrics.

MRR — Mean Reciprocal Rank

Find the rank of the first relevant document, take its reciprocal, average over queries. Rank 1 scores 1.0, rank 2 scores 0.5, rank 4 scores 0.25, not found scores 0.

MRR=1∣Q∣∑i=1∣Q∣1ranki\text{MRR} = \frac{1}{|Q|}\sum_{i=1}^{|Q|} \frac{1}{\text{rank}_i}

QueryRank beforeRR beforeRank afterRR after
111.000011.0000
230.333311.0000
320.500020.5000
4not found0.000040.2500
550.200011.0000
Sum2.03333.7500

MRR before: 2.0333/5=0.40672.0333 / 5 = 0.4067. After: 3.75/5=0.75003.75 / 5 = 0.7500 — an 84% relative gain purely from reordering documents the retriever had already found. Query 4 is the opening scenario, fixed. MRR is right when there is one correct answer; it is wrong when several documents are relevant, because it ignores everything after the first hit.

NDCG@K — Normalised Discounted Cumulative Gain

NDCG handles graded relevance (3 = perfect, 2 = useful, 1 = marginal, 0 = irrelevant) and rewards putting the best material highest. Gain at rank ii is discounted by log⁡2(i+1)\log_2(i+1):

DCG@K=∑i=1Krelilog⁡2(i+1)NDCG@K=DCG@KIDCG@K\text{DCG@K} = \sum_{i=1}^{K} \frac{rel_i}{\log_2(i+1)} \qquad \text{NDCG@K} = \frac{\text{DCG@K}}{\text{IDCG@K}}

Say the retriever returns grades [0, 2, 0, 3, 1] at ranks 1-5:

Rank iirelirel_ilog⁡2(i+1)\log_2(i+1)Contribution
101.00000.0000
221.58501.2619
302.00000.0000
432.32191.2921
512.58500.3868
DCG@52.9408

The ideal ordering of those same grades is [3, 2, 1, 0, 0]:

IDCG@5=31.0000+21.5850+12.0000=3.0000+1.2619+0.5000=4.7619\text{IDCG@5} = \frac{3}{1.0000} + \frac{2}{1.5850} + \frac{1}{2.0000} = 3.0000 + 1.2619 + 0.5000 = 4.7619

So NDCG@5=2.9408/4.7619=0.618\text{NDCG@5} = 2.9408 / 4.7619 = \mathbf{0.618}.

Now re-rank to [3, 2, 0, 1, 0]: DCG@5=3.0000+1.2619+0+12.3219+0=4.6926\text{DCG@5} = 3.0000 + 1.2619 + 0 + \frac{1}{2.3219} + 0 = 4.6926, giving NDCG@5=4.6926/4.7619=0.985\text{NDCG@5} = 4.6926 / 4.7619 = \mathbf{0.985}.

0.618 to 0.985 — same five documents, different order.

One trap: two gain formulas are in common use. The linear one above uses relirel_i; the exponential one uses 2reli−12^{rel_i} - 1, weighting a grade-3 document seven times a grade-1 rather than three times. On these same lists it gives 0.564 before and 0.993 after. Both are called "NDCG@5", so always state which you used and never compare across conventions.

The latency and accuracy trade-off

Every re-ranking decision spends milliseconds to buy precision. Whether that trade is good depends entirely on what the system is for.

ConfigurationAdded latencyNDCG@5 (illustrative)Fits
No re-ranking0 ms0.61Autocomplete, live suggestions
MiniLM-L-6, fetch_k=20+50 ms0.74Interactive chat — the default choice
MiniLM-L-6, fetch_k=50+120 ms0.83Interactive chat where accuracy matters
bge-reranker-large, fetch_k=50+320 ms0.88Support, research, internal tools
Cascade + LLM listwise+1,100 ms0.92Legal, medical, compliance; batch jobs

Budget against the whole request, not the re-ranker alone. If generation already takes 900 ms, a 120 ms re-ranker adds 13% to perceived latency and lifts NDCG by twenty-two points — not a close call. Serving typeahead in 40 ms, the same re-ranker is a 300% increase and obviously wrong.

Three levers to claw latency back, cheapest first: cut fetch_k (linear saving, flat curve above the knee); drop to a smaller re-ranker (MiniLM-L-2 is about half the cost of L-6); cache scores, which are deterministic per query-document pair.

Where this goes wrong

Six failure modes account for nearly everything that breaks here.

SymptomCauseFix
Re-ranking barely moves accuracyfetch_k too small — the answer was never in the candidate setMeasure recall@fetch_k first; raise fetch_k until it plateaus
Long chunks score inexplicably lowCross-encoders truncate at 512 tokens; the relevant sentence was cut offChunk under ~400 tokens, or score a summary/window of the chunk
Thresholds stop working after a model swapLogit scales are model-specific and uncalibratedRe-tune thresholds on a held-out set for every model change
Retrieval returns 2 docs where 50 were expectedMetadata filter applied after retrieval instead of inside the indexPush the predicate into the vector store's filtered search
Five chunks all say the same thingNo redundancy filter; corpus has versioned near-duplicatesMMR with λ≈0.7\lambda \approx 0.7, and deduplicate at ingestion
Confident wrong answers on out-of-scope questionsNo confidence gate; a flat score profile was treated as a good resultScore top + margin + support; refuse below threshold

The first is subtlest and worth naming plainly: re-ranking cannot create recall. If the top 50 contains the answer 68% of the time, 68% is your absolute accuracy ceiling however good the re-ranker is. Teams routinely spend a fortnight benchmarking re-rankers when the real problem is a chunker that split the answer across two chunks so neither holds it whole.

What this means when you build one

Start by measuring the ceiling. Take 50 real questions with known answers, retrieve at k=50, and record how often the answer appears anywhere in that list. That one number tells you which problem you have: below about 0.85 it is chunking, embeddings or missing keyword search, and a re-ranker will not help; above 0.85 it is ordering, and a re-ranker is the highest-return change available.

When it is an ordering problem, add the smallest thing that works: ms-marco-MiniLM-L-6-v2, fetch_k=50, top_k=5. Three lines and about 120 ms. Measure NDCG@5 and MRR before and after on the same 50 questions; you want a jump of ten points or more. If you do not get it, the ceiling was your problem after all.

Then add filters in order: metadata pre-filtering inside the index, a score threshold with a floor guaranteeing one survivor, and MMR at λ=0.7\lambda = 0.7 if your corpus has versioned documents — it almost certainly does.

Last, add the confidence gate and make it visible: log the score on every request alongside the query. Within a week you will have a ranked list of the questions your corpus genuinely cannot answer — the most actionable document a RAG team can own, because it tells you what to write, what to ingest, and what to stop pretending you cover. The compliance team above found their wire-transfer failure that way, not from a bug report but from a cluster of low-margin queries all circling the same missing authority matrix.