Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your knowledge base is in English but users ask in Hindi, Tamil and Hinglish, and retrieval collapses. What do you do?
What you need to know
An English-only model has never learned that mapping. It sees Devanagari characters as unfamiliar tokens, and the query vector ends up in a random-looking region. Nothing errors; the top results are simply wrong.
The two options
A: Multilingual embeddings
- One model for all languages (BGE-M3, multilingual-e5-large, Cohere's multilingual models, Qwen3-Embedding)
- No translation step at query time
- Full re-index of the corpus
- Can be slightly weaker than an English specialist on English-only queries
B: Translate the query
- Keep the English index untouched
- A fast model translates the query first
- Adds a few hundred milliseconds
- Weak on code-mixed Hinglish and on local product names
Choose A for a product that will be multilingual permanently. Choose B when the corpus is huge, re-indexing is expensive, and non-English traffic is a small slice. Many teams start with B and move to A.
Two details that get missed
- BM25 has no cross-lingual ability at all. It matches exact words. A Hindi query shares no words with an English chunk, so the lexical leg of hybrid search returns nothing useful. Feed it the translated query, and for Hinglish in Latin script, the transliterated and translated form.
- The answer language drifts. Given English context, models often answer in English. Say it explicitly: "Answer in the same language as the user's question."
1query_en = translate_to_english(user_query) # also used for the BM25 leg2dense = vector_index.search(embed(user_query), k=30) # multilingual model handles the original3sparse = bm25_index.search(query_en, k=30)4chunks = rerank(user_query, rrf_fuse(dense, sparse))[:5]5answer = llm.generate(prompt(chunks, user_query, reply_in=detect_language(user_query)))This combines both ideas: the multilingual dense leg uses the original query, the BM25 leg uses the English translation, and the reply language is set explicitly.
Measure per language
An aggregate score hides a broken language. If 85% of traffic is English at 0.90 recall and Tamil is at 0.20, the average still looks fine. Build a small eval set per language, even 50 questions each, and report every language separately.
A real-life example
Scenario (illustrative numbers). A state-run bank's help assistant has 12,000 English FAQ and policy chunks. After a marketing push in Tamil Nadu and Uttar Pradesh, 30% of questions arrive in Tamil, Hindi or Hinglish. Recall@5 on a per-language eval is 0.88 for English, 0.31 for Hindi, 0.22 for Tamil and 0.35 for Hinglish.
Re-embedding 12,000 chunks is cheap, so they switch to a multilingual model and re-index in under an hour. Hindi recall rises to 0.82 and Tamil to 0.81, while English holds at 0.87. Hinglish improves less, to 0.62, because users write "mera account block kyu hua" in Latin script. Adding a translate-and-transliterate step for the BM25 leg lifts Hinglish to 0.78. The prompt now sets the reply language, and complaints about "the bot replies in English" stop.
Follow-up questions to expect
- "Should you translate the whole corpus instead?" — Possible for a small, stable corpus, but you multiply storage and must re-translate on every update. Query-side approaches usually win.
- "How do you detect the language?" — A lightweight language-ID library, plus a script check (Devanagari, Tamil, Latin); Hinglish needs a small classifier or a model call because it uses Latin script.
- "Does the reranker need to be multilingual too?" — Yes. A multilingual cross-encoder, or rerank using the translated query.