Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
You just launched a RAG-powered product assistant. On Day 1, your vector DB has no embeddings yet but users are already asking questions. How do you handle the cold start problem so the system doesn't fail or hallucinate before the index is populated?
What you need to know
Why an empty index causes hallucination
A RAG prompt says "answer using this context". When the context is empty or irrelevant, many models still answer — from their training data, in the confident tone of a documented fact. So the danger on day one is not an error; it is a fluent wrong answer about your product.
1def answer(question: str):2 hits = hybrid_search(question, k=30) # works even with zero vectors3 ranked = rerank(question, hits)[:5]4 if not ranked or ranked[0].score < THRESHOLD:5 log_gap(question) # becomes tomorrow's ingestion list6 return {"text": "That isn't in my knowledge base yet. Here is our help centre, "7 "or I can connect you to a person.", "handoff": True}8 return generate_with_citations(question, ranked) # the prompt requires a citation per claimShrink the cold window
- Rank documents by demand — use the last 90 days of support tickets and site-search logs to estimate which docs answer the most questions, and embed those first.
- Serve keyword search now — BM25 or Postgres full-text search needs only raw text, so it works as soon as documents are loaded. Run hybrid search, and let vectors join as they are embedded.
- Seed FAQ pairs — fifty good question-answer pairs from the support team, indexed as high-priority chunks.
- Log every gated query — the gap list tells you what to ingest or write next.
| Approach | Ready | Quality on day one |
|---|---|---|
| Wait until everything is embedded | Hours to days | Good, but late |
| Keyword search only | Minutes | Fine for exact terms, weak for paraphrases |
| Demand-ordered embedding plus keyword | Minutes, improving hourly | Most common questions covered early |
Metrics
Coverage rate (share of questions with a chunk above the threshold), the "I don't know" rate falling day by day, and deflection rate (questions solved without a human). Require a citation for every claim, so a half-filled index still cannot bluff.
A real-life example
Scenario, numbers made up. A consumer-electronics brand launches a product assistant with 40,000 help articles and manuals. Embedding everything will take about 30 hours because of rate limits, and the launch email has already gone out.
The team loads raw text into Postgres full-text search in 20 minutes. Ticket logs show that about 2,000 articles (5%) matched roughly two-thirds of last quarter's questions, so those are embedded first, done in 90 minutes. Support writes 60 FAQ pairs about the new model. On day one, 34% of questions hit the coverage gate and get the honest fallback; by day three, 9%. The gap log shows 120 questions about a warranty-registration flow that had no article at all, and the support team writes one.
Follow-up questions to expect
- "How do you set the threshold with no data?" — Start conservative using a small hand-labelled set of 50–100 questions, then recalibrate on real traffic in the first week.
- "Why not let the model answer from general knowledge, with a disclaimer?" — For product facts, general knowledge is usually wrong or out of date, and users ignore disclaimers.
- "Is keyword search good enough on its own?" — For exact names and error codes, yes; for paraphrased questions, it misses. That is why vectors join the mix as they arrive.