Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your RAG system at HDFC Bank sometimes confidently gives wrong answers because the retrieved chunk was only partially relevant. How do you add a confidence/trust score to every RAG response so users know when to trust the output?


What you need to know

Why "partially relevant" is dangerous

A chunk about "home loan prepayment charges for floating rates" is partially relevant to a question about fixed-rate loans. The model sees a confident-looking source and writes a confident answer — for the wrong product. The score must notice that the evidence is weak even when the text sounds sure.

The signals

StageSignalWhy it helps
RetrievalTop-1 cross-encoder scoreMuch better relevance signal than raw cosine
RetrievalGap between rank 1 and rank 5A flat list means nothing stood out
RetrievalNumber of chunks above thresholdMany supporting chunks suggest solid evidence
GenerationNLI groundedness per sentenceCatches claims the chunks do not support
GenerationAgreement across 3 samplesDisagreement hints at guessing
GenerationToken log-probabilities, if availableA weaker extra signal

Combine and calibrate

Python
import numpy as npfrom sklearn.linear_model import LogisticRegressionfrom sklearn.calibration import calibration_curve# X: one row per answer [top1, gap_1_5, n_above, grounded_share, agreement]; y: 1 if correctmodel = LogisticRegression().fit(X_train, y_train)p = model.predict_proba(X_test)[:, 1]frac_correct, mean_pred = calibration_curve(y_test, p, n_bins=10)threshold = min(t for t in np.linspace(0.5, 0.99, 50)                if (y_test[p >= t].mean() if (p >= t).any() else 0) >= 0.95)print("auto-answer when p >=", round(threshold, 2), "coverage:", (p >= threshold).mean())

Logistic regression gives a probability, not a heuristic, and its weights show which signals matter. The threshold is picked for a precision target (here 95% correct among auto-answered questions), and you accept the coverage that comes with it.

Act on bands

  1. High — answer normally, with citations.
  2. Medium — answer, but show the source passage and say "please verify against the policy document".
  3. Low — do not answer; route to a human agent with the retrieved context attached.

For a bank, the threshold is a business and risk decision, not an ML one. Log the score with every response so you can later explain why the system answered.

A real-life example

Scenario, numbers made up. A large Indian bank's internal assistant helps branch staff with product rules. An audit of 400 answers finds 11% wrong, mostly from partially relevant chunks, such as NRI account rules applied to resident accounts.

The team labels those 400 answers and fits a logistic regression on six signals. At a threshold giving 96% precision, 72% of questions are auto-answered, 18% are answered with the source shown, and 10% go to the product desk. Over the next month, wrong auto-answers fall to about 3%, and staff start trusting the "high" band enough to stop double-checking every answer.

Follow-up questions to expect

  • "Why not just use cosine similarity as confidence?" — Cosine measures topic overlap, not whether the chunk answers the question, and its scale varies by model and query.
  • "How often do you recalibrate?" — Whenever the model, prompt, retriever or corpus changes, and on a schedule, such as monthly, with fresh labels.
  • "What do users see?" — Bands with clear actions, not a raw percentage; a number without a meaning invites misuse.