Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your RAG retrieves top-5 chunks, but the correct answer lives in chunk #12. Increasing top-K to 20 blows the context window. How do you make sure the right chunk reaches the LLM without flooding the context?
What you need to know
Chunk #12 was found. It was just ranked badly. So the fix is better ranking inside a wide candidate set, not a bigger window.
Retrieve wide, rerank, read narrow
- Retrieve wide — top 50 from dense search and top 50 from BM25.
- Fuse — reciprocal rank fusion (RRF) merges the two ranked lists by position, so their different score scales don't matter.
- Rerank — a cross-encoder scores all candidates against the query.
- Read narrow — the top 5 go to the model, strongest first or last, never buried in the middle.
candidates = rrf_fuse(dense.search(query, k=50), bm25.search(query, k=50))[:50]scores = reranker.predict([(query, c.text) for c in candidates]) # e.g. a bge-reranker modelcontext = [c for _, c in sorted(zip(scores, candidates), key=lambda x: -x[0])][:5]The cross-encoder catches relevance that cosine similarity misses: negation, the specific entity asked about, whether the chunk actually answers or merely mentions the topic.
Cost and latency
| Option | Added latency | Added tokens | Accuracy effect |
|---|---|---|---|
| top-k = 20 into the prompt | Longer prefill | about 4x more context | Diluted; the model may miss the key chunk |
| Rerank 50, keep 5 | Tens of ms (local GPU) to a few hundred ms (hosted) | None extra | Usually the best precision |
Supporting fixes
- Query rewriting or multi-query. Chunk #12 often ranks low because of vocabulary mismatch ("cancel my SIM" versus "service termination"). Generating two or three rephrasings widens recall.
- Ordering. Models attend best to the start and end of the context. Put the strongest chunks there.
The diagnostic that decides where to work
Measure recall@50: is the gold chunk anywhere in the candidate set? If yes, reranking will help. If no, reranking cannot save you; the problem is upstream in chunking, embeddings or missing hybrid search.
A real-life example
Scenario (illustrative numbers). A telecom operator's support bot answers questions about recharge plans and porting. On 200 labelled questions, the gold chunk is in the top 5 only 61% of the time, but in the top 50 for 94%.
The team adds BM25 with RRF and a cross-encoder reranker running on the same GPU as their embedding model. precision@5 rises from 0.61 to 0.87, and median added latency is 45 ms. Prompt size stays at 5 chunks, about 2,000 tokens. Questions like "can I keep my number if I move to Pune" now retrieve the porting-policy chunk that used to sit around rank 12.
Follow-up questions to expect
- "Why not just use a bigger embedding model?" — It may help recall, but a cross-encoder reading the pair together is a bigger jump in ranking quality for a small latency cost.
- "How many candidates should the reranker see?" — Enough that recall@N is high, commonly 20 to 100; beyond that, latency grows with little gain.
- "Can an LLM be the reranker?" — Yes, and it can be accurate, but it is slower and more expensive than a dedicated cross-encoder; use it for small, high-value sets.