Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your RAG retrieves top-5 chunks, but the correct answer lives in chunk #12. Increasing top-K to 20 blows the context window. How do you make sure the right chunk reaches the LLM without flooding the context?


Retrieve wide, rerank, read narrowDense top 50plus BM25 top 50Reciprocalrank fusionCross-encoderscores 50 pairsTop 5 to themodel, about2K tokensGold chunk in top 50 for 94% of questions, top 5 for 61%.
Chunk 12 was found all along; it only needed a model that reads the question and the chunk together.

What you need to know

Chunk #12 was found. It was just ranked badly. So the fix is better ranking inside a wide candidate set, not a bigger window.

Retrieve wide, rerank, read narrow

  1. Retrieve wide — top 50 from dense search and top 50 from BM25.
  2. Fuse — reciprocal rank fusion (RRF) merges the two ranked lists by position, so their different score scales don't matter.
  3. Rerank — a cross-encoder scores all candidates against the query.
  4. Read narrow — the top 5 go to the model, strongest first or last, never buried in the middle.
Python
candidates = rrf_fuse(dense.search(query, k=50), bm25.search(query, k=50))[:50]scores = reranker.predict([(query, c.text) for c in candidates])   # e.g. a bge-reranker modelcontext = [c for _, c in sorted(zip(scores, candidates), key=lambda x: -x[0])][:5]

The cross-encoder catches relevance that cosine similarity misses: negation, the specific entity asked about, whether the chunk actually answers or merely mentions the topic.

Cost and latency

OptionAdded latencyAdded tokensAccuracy effect
top-k = 20 into the promptLonger prefillabout 4x more contextDiluted; the model may miss the key chunk
Rerank 50, keep 5Tens of ms (local GPU) to a few hundred ms (hosted)None extraUsually the best precision

Supporting fixes

  • Query rewriting or multi-query. Chunk #12 often ranks low because of vocabulary mismatch ("cancel my SIM" versus "service termination"). Generating two or three rephrasings widens recall.
  • Ordering. Models attend best to the start and end of the context. Put the strongest chunks there.

The diagnostic that decides where to work

Measure recall@50: is the gold chunk anywhere in the candidate set? If yes, reranking will help. If no, reranking cannot save you; the problem is upstream in chunking, embeddings or missing hybrid search.

A real-life example

Scenario (illustrative numbers). A telecom operator's support bot answers questions about recharge plans and porting. On 200 labelled questions, the gold chunk is in the top 5 only 61% of the time, but in the top 50 for 94%.

The team adds BM25 with RRF and a cross-encoder reranker running on the same GPU as their embedding model. precision@5 rises from 0.61 to 0.87, and median added latency is 45 ms. Prompt size stays at 5 chunks, about 2,000 tokens. Questions like "can I keep my number if I move to Pune" now retrieve the porting-policy chunk that used to sit around rank 12.

Follow-up questions to expect

  • "Why not just use a bigger embedding model?" — It may help recall, but a cross-encoder reading the pair together is a bigger jump in ranking quality for a small latency cost.
  • "How many candidates should the reranker see?" — Enough that recall@N is high, commonly 20 to 100; beyond that, latency grows with little gain.
  • "Can an LLM be the reranker?" — Yes, and it can be accurate, but it is slower and more expensive than a dedicated cross-encoder; use it for small, high-value sets.