Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your RAG system retrieves top-20 chunks, but the most relevant chunk is always ranked #12 because cosine similarity alone doesn't capture semantic importance. How do you add a re-ranking layer to improve answer quality without blowing up latency?
What you need to know
Bi-encoder versus cross-encoder
Bi-encoder (embedding search)
- Embeds query and chunk separately
- Chunk vectors computed once, ahead of time
- Search over millions in milliseconds
- Scores "same topic", not "answers it"
Cross-encoder (reranker)
- Reads query and chunk together
- Must run for every pair at query time
- Too slow for millions; fine for 50
- Scores whether this chunk answers this question
Because the cross-encoder sees both texts at once, it can notice that a chunk mentions "refund" but talks about a different product, or that the chunk ranked 12th contains the exact answer.
The two-stage design
- Recall stage — vector search plus BM25, fused, top 50. Its only job is to get the right chunk somewhere in the list.
- Precision stage — rerank the 50 with a cross-encoder, keep 3–5.
- Generate — with the short, well-ordered context.
Options include self-hosted open rerankers (for example the BGE reranker family) or hosted rerank APIs (Cohere, Voyage and others).
1from sentence_transformers import CrossEncoder23reranker = CrossEncoder("BAAI/bge-reranker-v2-m3", max_length=512)45def retrieve(query: str, k_final: int = 4):6 candidates = hybrid_search(query, k=50) # recall stage7 scores = reranker.predict([(query, c.text) for c in candidates], batch_size=50)8 ranked = sorted(zip(candidates, scores), key=lambda x: -x[1])9 return [c for c, _ in ranked[:k_final]] # precision stageThe latency accounting
| Change | Effect |
|---|---|
| Add reranking of 50 pairs on a GPU | Adds tens of ms; more on CPU or over a network hop |
| Cut prompt from 20 chunks to 4 | Far fewer input tokens, so faster prefill and lower cost |
| Batch all pairs in one call | Avoids 50 separate round trips |
| Cache (query hash, chunk ID) scores | Repeated questions cost nothing |
| Skip reranking when top-1 is clearly ahead | Saves time on easy queries |
Measure three numbers: recall@50 of the first stage (your ceiling — a reranker cannot rescue a chunk that was never retrieved), nDCG@5 or MRR before and after reranking (ranking quality), and p95 rerank latency.
A real-life example
Scenario, numbers made up. An HR assistant for a 20,000-employee company sends the top 20 chunks to the model. The right policy paragraph is in the 20 but often around rank 10–12, and the model sometimes answers from a similar policy for a different grade.
The team retrieves 50 with hybrid search (recall@50 is 94% on a 300-question golden set), reranks with a self-hosted cross-encoder on a small GPU (about 40 ms p95 for 50 pairs), and keeps 4. MRR rises from 0.41 to 0.78. The prompt shrinks from about 9,000 to 2,000 tokens, so end-to-end p95 latency falls by roughly a second, and answer accuracy on the golden set rises from 76% to 88%.
Follow-up questions to expect
- "Why not rerank 200 candidates?" — Cost grows with each pair, and gains flatten once recall is high. Measure recall at 50, 100 and 200 and stop where it levels off.
- "Can an LLM be the reranker?" — Yes, listwise LLM reranking works well, but it is slower and pricier; use it for small candidate sets or offline.
- "What if recall@50 is low?" — Fix the first stage — hybrid search, better chunking, query rewriting — because reranking only reorders what it is given.