Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your medical assistant says with 100% confidence that drug X treats condition Y — but the source paper says the opposite. How do you calibrate LLM confidence so users know when to doubt the model?


From a fluent answer to an honest confidence bandSplit theanswer intosingle claimsNLI checkagainstthe sourceAgreementacross 5 samplesIsotonic fit onexpert labelsShow, or routeto a pharmacistA contradicted claim is blocked outright, not scored lower.
The model sounded sure on 94 percent of answers but was right on 78 percent; only calibration against labels closes that gap.

What you need to know

What calibration means

LLMs are trained to write fluent text, and fluent text sounds sure. When asked "how confident are you, 0–100?", models tend to be overconfident. So the model's words cannot be the number you show.

Three signals, strongest first

SignalHow it worksWeakness
Grounding checkSplit the answer into single claims; an NLI model (natural language inference) labels each as supported, neutral or contradicted by the sourceOnly as good as retrieval
Self-consistencySample 5 answers at moderate temperature; measure agreement on the key claimMisses answers that are consistently wrong
Token probabilitiesLow average probability on content tokens hints at guessingMany APIs, including most reasoning models, do not expose them
Verbalised confidenceAsk the model for a numberWeakest; use only as one input

A contradiction is not a lower score. It is a hard block: the answer says the opposite of its own source.

Calibrate the combined score

Combine the signals into one raw score, then fit it to real correctness on a labelled set with isotonic regression, which maps raw scores to observed accuracy.

Python
import numpy as npfrom sklearn.isotonic import IsotonicRegressiondef ece(conf, correct, bins=10):    conf, correct = np.asarray(conf), np.asarray(correct, float)    edges, total = np.linspace(0, 1, bins + 1), 0.0    for lo, hi in zip(edges[:-1], edges[1:]):        m = (conf > lo) & (conf <= hi)        if m.any():            total += m.mean() * abs(conf[m].mean() - correct[m].mean())    return total# raw = 0.5*grounding + 0.3*agreement + 0.2*token_prob   (on 400 expert-labelled answers)# iso = IsotonicRegression(out_of_bounds="clip").fit(raw_train, correct_train)# print(ece(iso.predict(raw_test), correct_test))

Check the result with a reliability diagram (shown confidence on one axis, actual accuracy on the other) and ECE on a held-out split. Recalibrate when the model, prompt or corpus changes.

The product surface

A number alone is not enough. Below the threshold, do not show a percentage; show the source passage, say that the answer is uncertain, and route to a clinician or pharmacist. Log every high-confidence error as a top-priority incident, because those are the ones that cause harm.

A real-life example

Scenario, numbers made up. A hospital group's drug-information assistant is tested on 400 questions labelled by pharmacists. The model sounds "highly confident" on 94% of answers, but only 78% are correct.

The team adds claim-level NLI against retrieved monographs, five-sample agreement and token probabilities from an open-weights model they host. After isotonic calibration on 300 answers and testing on 100, answers shown as "high confidence" are correct about 96% of the time, and 22% of answers are routed to "check the source". The drug X case is now blocked: the NLI check marks the claim "X treats Y" as contradicted by the retrieved paper.

Follow-up questions to expect

  • "Why not just ask the model how sure it is?" — Verbalised confidence is poorly calibrated, and it tracks tone more than correctness. It can be one weak input, never the displayed number.
  • "Does self-consistency catch everything?" — No. A model can be consistently wrong, which is why the grounding check against sources comes first.
  • "What if your API gives no token probabilities?" — Use grounding and self-consistency; they need only text, and they are usually the stronger signals anyway.