Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your RAG system at HDFC Bank sometimes confidently gives wrong answers because the retrieved chunk was only partially relevant. How do you add a confidence/trust score to every RAG response so users know when to trust the output?
What you need to know
Why "partially relevant" is dangerous
A chunk about "home loan prepayment charges for floating rates" is partially relevant to a question about fixed-rate loans. The model sees a confident-looking source and writes a confident answer — for the wrong product. The score must notice that the evidence is weak even when the text sounds sure.
The signals
| Stage | Signal | Why it helps |
|---|---|---|
| Retrieval | Top-1 cross-encoder score | Much better relevance signal than raw cosine |
| Retrieval | Gap between rank 1 and rank 5 | A flat list means nothing stood out |
| Retrieval | Number of chunks above threshold | Many supporting chunks suggest solid evidence |
| Generation | NLI groundedness per sentence | Catches claims the chunks do not support |
| Generation | Agreement across 3 samples | Disagreement hints at guessing |
| Generation | Token log-probabilities, if available | A weaker extra signal |
Combine and calibrate
1import numpy as np2from sklearn.linear_model import LogisticRegression3from sklearn.calibration import calibration_curve45# X: one row per answer [top1, gap_1_5, n_above, grounded_share, agreement]; y: 1 if correct6model = LogisticRegression().fit(X_train, y_train)7p = model.predict_proba(X_test)[:, 1]89frac_correct, mean_pred = calibration_curve(y_test, p, n_bins=10)10threshold = min(t for t in np.linspace(0.5, 0.99, 50)11 if (y_test[p >= t].mean() if (p >= t).any() else 0) >= 0.95)12print("auto-answer when p >=", round(threshold, 2), "coverage:", (p >= threshold).mean())Logistic regression gives a probability, not a heuristic, and its weights show which signals matter. The threshold is picked for a precision target (here 95% correct among auto-answered questions), and you accept the coverage that comes with it.
Act on bands
- High — answer normally, with citations.
- Medium — answer, but show the source passage and say "please verify against the policy document".
- Low — do not answer; route to a human agent with the retrieved context attached.
For a bank, the threshold is a business and risk decision, not an ML one. Log the score with every response so you can later explain why the system answered.
A real-life example
Scenario, numbers made up. A large Indian bank's internal assistant helps branch staff with product rules. An audit of 400 answers finds 11% wrong, mostly from partially relevant chunks, such as NRI account rules applied to resident accounts.
The team labels those 400 answers and fits a logistic regression on six signals. At a threshold giving 96% precision, 72% of questions are auto-answered, 18% are answered with the source shown, and 10% go to the product desk. Over the next month, wrong auto-answers fall to about 3%, and staff start trusting the "high" band enough to stop double-checking every answer.
Follow-up questions to expect
- "Why not just use cosine similarity as confidence?" — Cosine measures topic overlap, not whether the chunk answers the question, and its scale varies by model and query.
- "How often do you recalibrate?" — Whenever the model, prompt, retriever or corpus changes, and on a schedule, such as monthly, with fresh labels.
- "What do users see?" — Bands with clear actions, not a raw percentage; a number without a meaning invites misuse.