Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Scenario – 1: Retrieval Pipeline Failure


Scenario: a hand-built retrieval pipeline, plain Python with an embeddings API and NumPy or FAISS, starts returning irrelevant chunks right after a batch of new documents lands. How do you find and fix the problem?

Bisect a hand-built retriever at its seamsVectors: same model, dimension, prefixesMapping: matrix row i is chunk iText: the chunk itself is readableGolden set of 50 queries on every ingest
Every broken state still produces a confident ranking, so only these checks, not the scores, reveal the break.

What you need to know

A hand-built retriever has three parts that must agree: the vectors, the mapping from vector rows to chunks, and the text inside the chunks. When results suddenly go bad after an ingest, one of these broke. Nothing throws an error, because every broken state still produces a valid-looking ranking.

Seam 1: are the vectors right?

  • Dimension mismatch. Someone switched embedding models for the new batch. Old vectors are 1,536-dimensional, new ones 768, or the same size but from a different model. Vectors from two models live in different spaces and cannot be compared.
  • Asymmetric prefixes. Some embedding models expect different prefixes for queries and documents. E5 models expect query: and passage: ; some BGE models expect an instruction on queries. Drop the prefix on one side and recall falls sharply with no error.
  • Normalisation. If you use a dot product, both sides must be normalised the same way.

Seam 2: is the id mapping right?

A typical pure-Python setup keeps a NumPy matrix of vectors plus a parallel list of chunk metadata. If a code path appends to one and not the other, or deletes and reorders one, every result points at the wrong chunk while the scores look perfect.

Python
assert emb.shape[0] == len(chunks), "vector rows and chunk list out of sync"assert emb.shape[1] == EXPECTED_DIM, "embedding dimension changed"q = embed("query: " + question)sims = (emb @ q) / (np.linalg.norm(emb, axis=1) * np.linalg.norm(q))top = [(chunks[i]["id"], float(sims[i])) for i in np.argsort(-sims)[:5]]

The two asserts run on every load and catch more real incidents than any tuning. Storing the chunk id alongside each vector, rather than relying on position, removes the problem entirely.

Seam 3: is the text right?

Print the actual text of the top results for a failing query. Common findings: a PDF extractor returning empty strings or garbled columns, chunks that are only a repeated header and footer, or a table flattened into meaningless numbers. All of these embed to valid vectors.

The debugging order

  1. Check shapes and model ids — seconds, and catches the most common break.
  2. Check the mapping — row counts, and a spot check that chunk i's text re-embeds close to vector i.
  3. Read the chunks — for five failing queries.
  4. Fix forward — a 50-query golden set run on every ingest, plus per-request logs of retrieved ids and scores.

A real-life example

Scenario (illustrative numbers). A small startup's internal knowledge bot uses a NumPy matrix of 40,000 chunk vectors. After a new batch of 3,000 HR documents is added, answers become random. The golden set, which they did not have yet, would have caught it in seconds.

The engineer runs the checks. Shapes match (43,000 rows and 43,000 chunks) and the dimension is right. But a spot check shows chunk 41,200's text is about travel expenses while its vector is closest to leave-policy text. The ingest script had sorted the new chunks by file name before saving the metadata list, but embedded them in upload order, so the counts matched while almost every new row pointed at the wrong chunk. Storing each vector together with its chunk id fixes it for good. A 50-query golden test now runs on every ingest, and recall@5 on it is back to 0.86.

Follow-up questions to expect

  • "When would you move from NumPy to FAISS or a vector DB?" — When brute-force search gets slow (hundreds of thousands to millions of vectors) or you need filtering, persistence and concurrent updates.
  • "How do you detect an embedding-model change automatically?" — Store the model name and dimension with the index and refuse to mix; check both at load time.
  • "What goes in the per-request log?" — The query, its embedding model, retrieved chunk ids and scores, and the final prompt.