Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Scenario – 1: Retrieval Quality Degradation


Scenario: a LangChain RAG chain that worked well at launch now returns weakly relevant documents as the corpus has grown. How do you diagnose and fix it?

What you need to know

"It worked at launch" is the clue. A small corpus hides retrieval weaknesses; a large one exposes them. As documents are added, there are more near-duplicates, more product versions, and more exact terms (error codes, SKUs) that dense embeddings blur together.

Recall versus ranking

Symptom from the tracesProblemFix
The right document is not in the top 25 at allRecallHybrid search, query rewriting, metadata filters
The right document is at rank 12, but the prompt takes the top 5RankingCross-encoder reranking
Results mix old and new product versionsScopeMetadata filters on product, version, date

Fixing recall: hybrid retrieval

Python
from langchain_classic.retrievers import EnsembleRetrieverfrom langchain_community.retrievers import BM25Retrieverhybrid = EnsembleRetriever(    retrievers=[vectorstore.as_retriever(search_kwargs={"k": 25}),                BM25Retriever.from_documents(docs, k=25)],    weights=[0.6, 0.4])

In LangChain 1.x the classic retrievers live in the langchain-classic package, and BM25Retriever comes from langchain-community, which is being wound down in favour of standalone integration packages, so check its current home before pinning versions. EnsembleRetriever merges the two ranked lists with weighted reciprocal rank fusion. BM25 catches exact tokens such as ERR_4032 or SKU-88120 that embeddings place near every similar code.

Fixing ranking: rerank

Python
from langchain_classic.retrievers import ContextualCompressionRetrieverfrom langchain_classic.retrievers.document_compressors import CrossEncoderRerankerfrom langchain_community.cross_encoders import HuggingFaceCrossEncoderreranker = CrossEncoderReranker(model=HuggingFaceCrossEncoder(model_name="BAAI/bge-reranker-base"), top_n=5)retriever = ContextualCompressionRetriever(base_compressor=reranker, base_retriever=hybrid)

The cross-encoder reads the query and each candidate together and re-sorts them; the top 5 go to the prompt.

Supporting fixes

  • Metadata filters. Pass search_kwargs={"k": 25, "filter": {"product": "v3"}} so a growing corpus doesn't put five product generations in one candidate pool. This is the most common cause of "it degraded as we grew".
  • MultiQueryRetriever generates paraphrases of the question when users phrase things inconsistently, and merges results.

The measurement loop

  1. Sample — 50 failing queries from LangSmith traces.
  2. Label — the correct document for each; save as a LangSmith dataset.
  3. Score — recall@25 for the retriever, precision@5 after reranking.
  4. Change one thing — rerun the dataset; keep the change only if the numbers improve.

Reranking 25 to 50 candidates typically adds tens of milliseconds on a local GPU, or a few hundred with a hosted API.

A real-life example

Scenario (illustrative numbers). A developer-tools company's docs assistant covered 1,200 pages at launch; a year later it covers 9,000 across three major product versions. Thumbs-down rate has doubled. The 50 sampled failures show 28 recall failures, mostly error-code queries and v2 pages returned for v3 questions, and 22 ranking failures.

The team adds BM25, a version filter taken from the user's workspace, and a cross-encoder reranker. On a 200-query LangSmith dataset, recall@25 rises from 0.71 to 0.93 and precision@5 from 0.52 to 0.81. Added latency is 90 ms at p50, and the thumbs-down rate returns to launch levels.

Follow-up questions to expect

  • "Is the in-memory BM25Retriever fine for production?" — For tens of thousands of documents, yes; beyond that, use the vector database's native hybrid or sparse search so it scales and persists.
  • "How do you choose the ensemble weights?" — Start at 50/50 and tune on the labelled dataset; keyword-heavy corpora often favour the sparse side.
  • "Why not just switch the embedding model?" — It might help recall slightly, but it costs a full re-index and doesn't fix exact-term or version-mixing problems.