Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your medical assistant says with 100% confidence that drug X treats condition Y — but the source paper says the opposite. How do you calibrate LLM confidence so users know when to doubt the model?
What you need to know
What calibration means
LLMs are trained to write fluent text, and fluent text sounds sure. When asked "how confident are you, 0–100?", models tend to be overconfident. So the model's words cannot be the number you show.
Three signals, strongest first
| Signal | How it works | Weakness |
|---|---|---|
| Grounding check | Split the answer into single claims; an NLI model (natural language inference) labels each as supported, neutral or contradicted by the source | Only as good as retrieval |
| Self-consistency | Sample 5 answers at moderate temperature; measure agreement on the key claim | Misses answers that are consistently wrong |
| Token probabilities | Low average probability on content tokens hints at guessing | Many APIs, including most reasoning models, do not expose them |
| Verbalised confidence | Ask the model for a number | Weakest; use only as one input |
A contradiction is not a lower score. It is a hard block: the answer says the opposite of its own source.
Calibrate the combined score
Combine the signals into one raw score, then fit it to real correctness on a labelled set with isotonic regression, which maps raw scores to observed accuracy.
1import numpy as np2from sklearn.isotonic import IsotonicRegression34def ece(conf, correct, bins=10):5 conf, correct = np.asarray(conf), np.asarray(correct, float)6 edges, total = np.linspace(0, 1, bins + 1), 0.07 for lo, hi in zip(edges[:-1], edges[1:]):8 m = (conf > lo) & (conf <= hi)9 if m.any():10 total += m.mean() * abs(conf[m].mean() - correct[m].mean())11 return total1213# raw = 0.5*grounding + 0.3*agreement + 0.2*token_prob (on 400 expert-labelled answers)14# iso = IsotonicRegression(out_of_bounds="clip").fit(raw_train, correct_train)15# print(ece(iso.predict(raw_test), correct_test))Check the result with a reliability diagram (shown confidence on one axis, actual accuracy on the other) and ECE on a held-out split. Recalibrate when the model, prompt or corpus changes.
The product surface
A number alone is not enough. Below the threshold, do not show a percentage; show the source passage, say that the answer is uncertain, and route to a clinician or pharmacist. Log every high-confidence error as a top-priority incident, because those are the ones that cause harm.
A real-life example
Scenario, numbers made up. A hospital group's drug-information assistant is tested on 400 questions labelled by pharmacists. The model sounds "highly confident" on 94% of answers, but only 78% are correct.
The team adds claim-level NLI against retrieved monographs, five-sample agreement and token probabilities from an open-weights model they host. After isotonic calibration on 300 answers and testing on 100, answers shown as "high confidence" are correct about 96% of the time, and 22% of answers are routed to "check the source". The drug X case is now blocked: the NLI check marks the claim "X treats Y" as contradicted by the retrieved paper.
Follow-up questions to expect
- "Why not just ask the model how sure it is?" — Verbalised confidence is poorly calibrated, and it tracks tone more than correctness. It can be one weak input, never the displayed number.
- "Does self-consistency catch everything?" — No. A model can be consistently wrong, which is why the grounding check against sources comes first.
- "What if your API gives no token probabilities?" — Use grounding and self-consistency; they need only text, and they are usually the stronger signals anyway.