Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your RAG system retrieves top-20 chunks, but the most relevant chunk is always ranked #12 because cosine similarity alone doesn't capture semantic importance. How do you add a re-ranking layer to improve answer quality without blowing up latency?


Two stages: recall, then precisionHybrid search:50 candidatesCross-encoderscores each pairKeep the top 4Prompt of 2,000tokens, not 9,000Recall at 50 was 94 percent — the ceiling the reranker works under.
Reranking added about 40 milliseconds but cut prefill so much that end-to-end latency fell by a second.

What you need to know

Bi-encoder versus cross-encoder

Bi-encoder (embedding search)

  • Embeds query and chunk separately
  • Chunk vectors computed once, ahead of time
  • Search over millions in milliseconds
  • Scores "same topic", not "answers it"

Cross-encoder (reranker)

  • Reads query and chunk together
  • Must run for every pair at query time
  • Too slow for millions; fine for 50
  • Scores whether this chunk answers this question

Because the cross-encoder sees both texts at once, it can notice that a chunk mentions "refund" but talks about a different product, or that the chunk ranked 12th contains the exact answer.

The two-stage design

  1. Recall stage — vector search plus BM25, fused, top 50. Its only job is to get the right chunk somewhere in the list.
  2. Precision stage — rerank the 50 with a cross-encoder, keep 3–5.
  3. Generate — with the short, well-ordered context.

Options include self-hosted open rerankers (for example the BGE reranker family) or hosted rerank APIs (Cohere, Voyage and others).

Python
from sentence_transformers import CrossEncoderreranker = CrossEncoder("BAAI/bge-reranker-v2-m3", max_length=512)def retrieve(query: str, k_final: int = 4):    candidates = hybrid_search(query, k=50)                        # recall stage    scores = reranker.predict([(query, c.text) for c in candidates], batch_size=50)    ranked = sorted(zip(candidates, scores), key=lambda x: -x[1])    return [c for c, _ in ranked[:k_final]]                       # precision stage

The latency accounting

ChangeEffect
Add reranking of 50 pairs on a GPUAdds tens of ms; more on CPU or over a network hop
Cut prompt from 20 chunks to 4Far fewer input tokens, so faster prefill and lower cost
Batch all pairs in one callAvoids 50 separate round trips
Cache (query hash, chunk ID) scoresRepeated questions cost nothing
Skip reranking when top-1 is clearly aheadSaves time on easy queries

Measure three numbers: recall@50 of the first stage (your ceiling — a reranker cannot rescue a chunk that was never retrieved), nDCG@5 or MRR before and after reranking (ranking quality), and p95 rerank latency.

A real-life example

Scenario, numbers made up. An HR assistant for a 20,000-employee company sends the top 20 chunks to the model. The right policy paragraph is in the 20 but often around rank 10–12, and the model sometimes answers from a similar policy for a different grade.

The team retrieves 50 with hybrid search (recall@50 is 94% on a 300-question golden set), reranks with a self-hosted cross-encoder on a small GPU (about 40 ms p95 for 50 pairs), and keeps 4. MRR rises from 0.41 to 0.78. The prompt shrinks from about 9,000 to 2,000 tokens, so end-to-end p95 latency falls by roughly a second, and answer accuracy on the golden set rises from 76% to 88%.

Follow-up questions to expect

  • "Why not rerank 200 candidates?" — Cost grows with each pair, and gains flatten once recall is high. Measure recall at 50, 100 and 200 and stop where it levels off.
  • "Can an LLM be the reranker?" — Yes, listwise LLM reranking works well, but it is slower and pricier; use it for small candidate sets or offline.
  • "What if recall@50 is low?" — Fix the first stage — hybrid search, better chunking, query rewriting — because reranking only reorders what it is given.