Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your document ingestion pipeline receives 500K files daily. Embedding every file immediately is too expensive. How do you prioritize, batch, and schedule embeddings efficiently at scale?


What survives the gate, and when it is embeddedrecent orsearchedwithin minutesdense and BM25normal priorityhourly batchdense and BM25rarely queriedon first accessBM25 until thenGoes here ifEmbeddedFindable byHotWarmColdDuplicates, near-duplicates and log exports are removed before any lane.
Keeping keyword search over everything is what makes it safe to embed the cold lane late, or never.

What you need to know

Gate first, because the cheapest embedding is the one you skip

  1. Exact dedupe — a SHA-256 hash of the normalised content; the same file uploaded to five folders is embedded once.
  2. Near-duplicate detection — MinHash on word shingles finds files that are 95% the same, such as a report with a new date.
  3. Chunk-level diff — for an edited file, hash each chunk and re-embed only the changed chunks.
  4. Type and size filters — build artefacts, logs and auto-generated exports rarely answer anyone's question.
  5. Score and route — send what is left to a lane.
Python
from datasketch import MinHash, MinHashLSHlsh = MinHashLSH(threshold=0.9, num_perm=128)def is_near_duplicate(doc_id: str, text: str) -> bool:    m = MinHash(num_perm=128)    words = text.lower().split()    for shingle in {" ".join(words[i:i + 5]) for i in range(len(words) - 4)}:        m.update(shingle.encode("utf8"))    if lsh.query(m):        return True    lsh.insert(doc_id, m)    return False

MinHash turns a document into a short signature, and locality-sensitive hashing (LSH) finds signatures with estimated similarity above 0.9 without comparing every pair.

Lanes

LaneWhat goes inWhen embeddedSearchable by
HotRecent, high-traffic source, or asked for by a searchWithin minutes, streamingDense and BM25
WarmNormal priorityHourly batchDense and BM25
ColdRarely queried folders, old archivesOn first access, or overnight in a batch jobBM25 until embedded

The value score combines recency, source importance, team, and the historical query hit rate for that folder or space. Start with a hand-written formula; a learned model is not the bottleneck. For cold work, provider batch APIs (OpenAI and Anthropic both offer about half price for asynchronous jobs), or spot GPUs for a self-hosted model, cut cost further.

Batching well

Group inputs into the batch size your model serves best, and sort by token length, so a batch does not pad 50 short texts up to the length of one long one.

Guard the backlog

Alert on queue depth and on the age of the oldest item per lane. When the hot lane falls behind, move lower-score items to cold instead of dropping them.

Measure cost per 1,000 documents, lag per lane, and the share of queries answered from hot-lane content.

A real-life example

Scenario, numbers made up. A bank's document platform receives 500,000 files a day and embeds everything as it arrives, costing about ₹9 lakh a month in embedding calls and falling 14 hours behind at month-end.

Gating removes 18% as exact duplicates, 12% as near-duplicates and 25% as log and export files nobody searches. Chunk-level diffs cut the remaining re-embedding of edited files by two thirds. Of what is left, 30,000 files a day go hot, 90,000 warm, and the rest cold with BM25 only. Monthly embedding cost falls to about ₹2.5 lakh, hot-lane lag stays under 5 minutes at month-end, and 93% of answers cite hot or warm content.

Follow-up questions to expect

  • "What if someone needs a cold document right now?" — BM25 finds it immediately, and the first access pushes it into the hot lane.
  • "Self-host the embedding model or use an API?" — At this volume, a self-hosted open embedding model on GPUs is often cheaper per token, but it adds operations work; compare on your own numbers.
  • "How do you tune the value score?" — Log which documents searches actually use; if many answers come from cold content, those folders deserve a higher score.