Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your document ingestion pipeline receives 500K files daily. Embedding every file immediately is too expensive. How do you prioritize, batch, and schedule embeddings efficiently at scale?
What you need to know
Gate first, because the cheapest embedding is the one you skip
- Exact dedupe — a SHA-256 hash of the normalised content; the same file uploaded to five folders is embedded once.
- Near-duplicate detection — MinHash on word shingles finds files that are 95% the same, such as a report with a new date.
- Chunk-level diff — for an edited file, hash each chunk and re-embed only the changed chunks.
- Type and size filters — build artefacts, logs and auto-generated exports rarely answer anyone's question.
- Score and route — send what is left to a lane.
1from datasketch import MinHash, MinHashLSH23lsh = MinHashLSH(threshold=0.9, num_perm=128)45def is_near_duplicate(doc_id: str, text: str) -> bool:6 m = MinHash(num_perm=128)7 words = text.lower().split()8 for shingle in {" ".join(words[i:i + 5]) for i in range(len(words) - 4)}:9 m.update(shingle.encode("utf8"))10 if lsh.query(m):11 return True12 lsh.insert(doc_id, m)13 return FalseMinHash turns a document into a short signature, and locality-sensitive hashing (LSH) finds signatures with estimated similarity above 0.9 without comparing every pair.
Lanes
| Lane | What goes in | When embedded | Searchable by |
|---|---|---|---|
| Hot | Recent, high-traffic source, or asked for by a search | Within minutes, streaming | Dense and BM25 |
| Warm | Normal priority | Hourly batch | Dense and BM25 |
| Cold | Rarely queried folders, old archives | On first access, or overnight in a batch job | BM25 until embedded |
The value score combines recency, source importance, team, and the historical query hit rate for that folder or space. Start with a hand-written formula; a learned model is not the bottleneck. For cold work, provider batch APIs (OpenAI and Anthropic both offer about half price for asynchronous jobs), or spot GPUs for a self-hosted model, cut cost further.
Batching well
Group inputs into the batch size your model serves best, and sort by token length, so a batch does not pad 50 short texts up to the length of one long one.
Guard the backlog
Alert on queue depth and on the age of the oldest item per lane. When the hot lane falls behind, move lower-score items to cold instead of dropping them.
Measure cost per 1,000 documents, lag per lane, and the share of queries answered from hot-lane content.
A real-life example
Scenario, numbers made up. A bank's document platform receives 500,000 files a day and embeds everything as it arrives, costing about ₹9 lakh a month in embedding calls and falling 14 hours behind at month-end.
Gating removes 18% as exact duplicates, 12% as near-duplicates and 25% as log and export files nobody searches. Chunk-level diffs cut the remaining re-embedding of edited files by two thirds. Of what is left, 30,000 files a day go hot, 90,000 warm, and the rest cold with BM25 only. Monthly embedding cost falls to about ₹2.5 lakh, hot-lane lag stays under 5 minutes at month-end, and 93% of answers cite hot or warm content.
Follow-up questions to expect
- "What if someone needs a cold document right now?" — BM25 finds it immediately, and the first access pushes it into the hot lane.
- "Self-host the embedding model or use an API?" — At this volume, a self-hosted open embedding model on GPUs is often cheaper per token, but it adds operations work; compare on your own numbers.
- "How do you tune the value score?" — Log which documents searches actually use; if many answers come from cold content, those folders deserve a higher score.