Medical AI is an evaluation problem, not a modelling problem

JR

Jai Rao

August 22, 202617 min read

Sensitivity, specificity, and the predictive-value arithmetic that explains why a 95% accurate screening model can still be wrong most times it fires.


Train a reasonable convolutional network on a public chest X-ray archive and you can post a strong AUC by the end of the afternoon. Modelling has not been the bottleneck in medical AI for years. The bottleneck is a question that sounds trivial and isn't: when this thing fires, what should the person on the other end of it actually believe?

That question is what separates clinical machine learning from most other applied ML. In consumer software a wrong prediction costs somebody a click. In a hospital it costs a needle biopsy, or six months of undetected disease progression, or an hour of the only radiologist on shift. The numbers that govern how often each of those happens are not the numbers most teams optimise, and they behave in ways that surprise people the first time they work them out on paper. What follows is the evaluation machinery — sensitivity, specificity, predictive value at realistic prevalence, and the quiet ways a model stops working when it changes buildings — because that machinery, not the architecture, is where these projects are won and lost.

Two numbers that pull in opposite directions

A clinical classifier almost always produces a continuous score, and then somebody picks a cutoff. Everything downstream depends on where that cutoff lands, so the two metrics that matter are both properties of the score plus the cutoff, never of the model alone.

Sensitivity is the fraction of genuinely sick people the model flags: true positives divided by all the people who actually have the condition. It is the metric that counts misses. Specificity is the fraction of genuinely healthy people the model clears: true negatives divided by all the people who actually don't have the condition. It is the metric that counts false alarms.

Both are conditioned on the truth rather than on the prediction, which has one very useful consequence: they don't change when the disease becomes more or less common. Move a 90%-sensitive test from a screening clinic to an emergency department and, image quality aside, it stays roughly 90% sensitive. That portability is why vendors quote them, and it's genuinely the right thing to quote. It is also the reason they answer a question no clinician is asking, which we'll get to shortly.

The two numbers are tied together by the threshold. Lower the cutoff and you catch more disease and raise more alarms; raise it and you do the reverse. The code below takes ten scored cases with known labels and reads sensitivity and specificity at three different cutoffs — watch how far the pair moves without a single weight changing.

Text
scores = [0.02, 0.11, 0.18, 0.31, 0.44, 0.52, 0.61, 0.73, 0.85, 0.94]truth  = [0,    0,    1,    0,    0,    1,    0,    1,    1,    1   ]def sens_spec(threshold):    tp = sum(1 for s, y in zip(scores, truth) if s >= threshold and y == 1)    fn = sum(1 for s, y in zip(scores, truth) if s <  threshold and y == 1)    fp = sum(1 for s, y in zip(scores, truth) if s >= threshold and y == 0)    tn = sum(1 for s, y in zip(scores, truth) if s <  threshold and y == 0)    return tp / (tp + fn), tn / (tn + fp)for t in (0.15, 0.50, 0.80):    se, sp = sens_spec(t)    print(f"threshold {t:.2f}  sensitivity {se:.0%}  specificity {sp:.0%}")

The output is 100% / 40%, then 80% / 80%, then 40% / 100%. One trained model, three completely different products: a net that catches everything and buries the reader in noise, a balanced compromise, and a tool that never cries wolf but walks past three cases in five. Nothing in the training run tells you which of those three you should ship.

The exchange rate between a miss and a false alarm

Picking the threshold means pricing two harms against each other, and they are not remotely the same shape.

A missed cancer on a screening study is not a neutral event that gets corrected next week. The disease keeps growing until the next screening round or until symptoms arrive, and for many cancers the stage at diagnosis is the single strongest lever on survival. A miss can quietly convert a treatable disease into an untreatable one, and the patient has been actively reassured in the meantime, which can make them slower to come back when something feels wrong.

A false alarm is not neutral either, and engineers routinely underrate it. It means a recall visit, more imaging, often a biopsy with its own small risk of bleeding and infection, and weeks of fear in between. In some screening contexts it means overdiagnosis: finding and treating a lesion that would never have harmed the person, with all the morbidity of treatment and none of the benefit. It also consumes exactly the scarce specialist capacity the model was supposed to create. Ten thousand extra recalls is not a rounding error in a health system; it is a staffing crisis.

There is no unit that converts one missed cancer into N unnecessary biopsies. That conversion is a value judgement about harm, and it belongs to clinicians, screening programme designers, patients, and sometimes regulators. The engineering obligation is narrower and quite firm: publish the full operating curve, make the threshold a deployment-time configuration rather than a constant in the training script, keep the score calibrated so that a stated probability means something stable, and never pick an operating point silently. A default cutoff of 0.5 inherited from a tutorial is picking one silently.

Why a 95% accurate screener is wrong most of the time it fires

Sensitivity and specificity are conditioned on the truth. The clinician's situation is the reverse: there is a positive result sitting in front of them and the truth is unknown. What they need is the probability of disease given a positive result — the positive predictive value — and its mirror, the probability of no disease given a negative result, the negative predictive value. Unlike sensitivity and specificity, these depend heavily on how common the disease is in the population being tested.

Work it through with round numbers. Take a model with 95% sensitivity and 95% specificity, applied to a screening population where the condition has a prevalence of 5 per 1,000. Screen 100,000 people.

  • 500 people have the disease; 99,500 do not.
  • Of the 500, the model flags 95% correctly: 475 true positives, and 25 missed.
  • Of the 99,500, the model clears 95%: 94,525 true negatives, and 4,975 false positives.

Overall accuracy is (475 + 94,525) / 100,000 = 95.0%, exactly as advertised. But the model raised 475 + 4,975 = 5,450 alarms, and only 475 of them were real. Positive predictive value is 475 / 5,450 = 8.7%. Roughly eleven out of every twelve times this "95% accurate" model fires, nobody is sick. Negative predictive value, meanwhile, is 94,525 / 94,550 = 99.97%.

The reason is structural rather than a flaw in this particular model. False positives are drawn from the enormous healthy majority, true positives from the tiny sick minority. Five percent of 99,500 is ten times larger than one hundred percent of 500. When disease is rare, the pool of alarms is dominated by healthy people, and improving sensitivity cannot fix it — even a perfect 100% sensitivity here only lifts PPV to 500 / 5,475, about 9.1%.

Here is the same arithmetic as code, swept across prevalences so the shape is visible rather than asserted.

Text
def screening_table(sensitivity, specificity, prevalences, n=100_000):    for p in prevalences:        disease, healthy = n * p, n * (1 - p)        tp, fn = sensitivity * disease, (1 - sensitivity) * disease        tn, fp = specificity * healthy, (1 - specificity) * healthy        ppv = tp / (tp + fp)        npv = tn / (tn + fn)        print(f"prev={p:7.3%}  TP={tp:8.1f}  FP={fp:8.1f}  PPV={ppv:7.2%}  NPV={npv:8.3%}")screening_table(0.95, 0.95, [0.001, 0.005, 0.01, 0.05, 0.10])

PPV climbs from 1.87% at a prevalence of 1 in 1,000 to 8.72% at 5 in 1,000, 16.10% at 1%, exactly 50.00% at 5%, and 67.86% at 10%. The model is unchanged throughout. Only the population moved.

PrevalenceTrue positivesFalse positivesPPVNPV
0.1%954,9951.9%99.995%
0.5%4754,9758.7%99.974%
1%9504,95016.1%99.947%
5%4,7504,75050.0%99.72%
10%9,5004,50067.9%99.42%

Notice the other half of the table. That same weak alarm is an excellent rule-out: a negative result at low prevalence is trustworthy to three or four decimal places. The honest reading is that this model is a poor diagnoser and a good triage filter, and a team that understood the arithmetic would build the product around the negatives instead of the positives.

What a trustworthy alarm actually costs

For a 95/95 model, PPV reaches 50% precisely when prevalence hits 5% — that's where sensitivity times prevalence equals the false-positive rate times its complement. Below that, the majority of alarms are false by construction. So if you want a defensible alarm at a realistic screening prevalence, you have to move something, and there are only three things to move.

The first is specificity, and the price is steep. This function inverts the PPV formula to ask what specificity a target PPV demands.

Text
def required_specificity(sensitivity, prevalence, target_ppv):    fpr = sensitivity * prevalence * (1 - target_ppv) / (target_ppv * (1 - prevalence))    return 1 - fprfor target in (0.10, 0.30, 0.50, 0.80):    spec = required_specificity(0.95, 0.005, target)    print(f"PPV {target:.0%} needs specificity {spec:.3%}  (false-positive rate {1 - spec:.3%})")

At 0.5% prevalence, a PPV of 50% requires 99.52% specificity — cutting the false-positive rate from 5% to under half a percent, a tenfold reduction, while holding sensitivity fixed. Pushing to 80% PPV needs 99.88%. That is usually not a hyperparameter search; it is a different sensing modality or a different problem.

The second lever is enrichment: change who gets tested. The same model applied to symptomatic referrals at 10% prevalence delivers a 68% PPV, and the arithmetic did all the work. The third is staging, which is what screening programmes have always done — a high-sensitivity, cheap first pass, then a specific confirmatory test that only ever sees the already-enriched positive pool. Both levers say the same thing: the model is one stage in a pipeline, and its metrics are meaningless without naming the population that feeds it.

One statistical trap deserves calling out here. At 0.5% prevalence, a 1,000-case test set contains about five positive cases. Any sensitivity estimated from it is close to noise, and the confidence interval will be embarrassing if anyone computes it. The rare class, not the total row count, is what bounds the strength of your evaluation, and gathering enough positives is frequently the most expensive part of the whole project.

The model that learned the scanner, not the disease

The second recurring failure has nothing to do with arithmetic. Models degrade when the data they meet stops resembling the data they were fitted on, and medical data shifts constantly.

Site-to-site covariate shift is the classic case. A model trained on images from one manufacturer's detector encounters another vendor's reconstruction kernel, a different dose protocol, different post-processing, a different patient mix and a different set of technologists positioning them. Accuracy drops. The dangerous part is not the drop but its silence: the network's confidence scores usually stay just as high, because nothing in the objective taught it to recognise inputs it has no business judging.

Shortcuts that look like skill

Worse than shift is a model that never learned the pathology in the first place. Networks are relentless at finding features correlated with the label in the training set, whether or not those features have anything to do with the disease. The pattern is uncomfortably consistent: models keying on burned-in text and laterality markers; on the signature of the portable scanner used at the bedside, which correlates with being sick enough not to travel to the imaging suite; on a chest drain visible in images of a pneumothorax that has already been treated; on the ruler marks and gel artefacts that appear beside skin lesions a clinician already thought worth measuring.

Every one of those shortcuts produces excellent held-out numbers and zero clinical value, because the correlation is an artefact of how the dataset was assembled and it will not follow the model to a new hospital. Saliency maps are weak evidence here — they are easy to read charitably. Occlusion and masking tests, training on deliberately cropped inputs, and checking whether a model can predict the source site from the image itself are all stronger. But the single most informative experiment is dull and non-negotiable: evaluate on data from a site the model has never seen, ideally with different equipment. Teams that skip external validation are not measuring generalisation; they are measuring their own dataset.

Unequal performance and the subgroups nobody measured

Training sets inherit whoever happened to walk through the doors of the contributing institutions. Dermatology image collections have historically skewed toward lighter skin tones, so a lesion classifier's learned features carry no guarantee of transferring across the full range of presentations. Imaging cohorts skew by age, sex, body habitus, comorbidity burden and geography. Physiological signals collected by devices with known measurement biases pass those biases straight into any model trained on them.

The structural problem is that aggregate metrics hide this by design. Overall sensitivity is a size-weighted average, and the group that was under-represented in training is usually under-represented in the test set too, so its poor performance barely moves the headline number. A model can be materially less sensitive for one group while the top-line AUC looks pristine.

The response is unglamorous. Pre-register the subgroups you will report before looking at results, publish per-subgroup sensitivity and specificity with confidence intervals, count how many positive cases sit in the smallest subgroup, and treat the widest interval as the honest bound on what you know. When a subgroup has eleven positives, you have no estimate — you have a placeholder. Fixing it generally requires collecting data rather than reweighting the loss.

There is a subtler version worth watching for. When the ground-truth labels were derived from historical clinical decisions, and those decisions were themselves distributed unevenly, the model learns the pattern of decisions rather than the pattern of disease. Evaluating against the same label source will never surface it, because the label and the bias come from the same place. Catching that needs a different definition of truth — biopsy results, registry follow-up, outcomes at twelve months — not a better metric.

Where AI is genuinely earning its place in the clinic

All of the above sounds discouraging, and it shouldn't be. It just explains why the deployments that work share a shape: they help a professional do something faster, and a professional still decides.

Triage and worklist prioritisation is the most deployable pattern in the field. The model doesn't diagnose anything; it reorders the reading queue so studies with a suspected time-critical finding surface first. The metric that matters is time-to-treatment for urgent cases, and the failure mode is bounded — a case the model ranks wrongly still gets read, just in its original place. That bounded downside is precisely why the pattern ships.

Documentation and note drafting has the largest immediate appetite, because clinician time lost to the keyboard is a real and measurable burden. Ambient capture drafting a visit note, a first pass at a discharge summary, coding assistance: the clinician reviews and signs, and the accountability never moves. The hazard is specific to this class — an error that survives review becomes part of a permanent record other people will rely on, and a draft that is fluent enough invites skimming rather than reading.

Imaging as a second reader fits programmes that already use double reading, where the second human is the scarce resource. The model reads independently, disagreements go to arbitration, and the comparison that matters is clinician-plus-model against clinician-alone.

Retinal and dermatological screening works for a set of converging reasons: a single organ, standardised image capture, well-defined grading scales, and a global shortage of the specialists who would otherwise look. Where specialist review is the bottleneck, a model that safely clears the unambiguous negatives creates capacity — which, as the NPV column showed, is what these models are actually good at.

One design risk cuts across all of them. A second reader that is right most of the time trains the human to stop reading independently. If the model's output is shown before the clinician has formed their own impression, you no longer have two readers; you have one reader and an anchor. The order in which the interface reveals things is a clinical safety decision, not a UX preference. Track override and disagreement rates in production: a system nobody ever overrides is either flawless or has become invisible, and only one of those is likely.

Retrospective accuracy is a hypothesis, not evidence

There is a ladder of evidence here, and the rungs are frequently conflated in vendor material and in engineering plans alike.

A retrospective result on a curated dataset with fixed labels answers one narrow question: could a model have separated these particular cases? It is necessary and it is cheap, and it is where nearly all published numbers live. A prospective evaluation runs the model on the real incoming stream, at real prevalence, with real image quality and real interruptions. Numbers commonly fall, sometimes considerably, and that fall is information rather than failure. An outcome study asks whether anything actually changed for patients — time to intervention, stage at diagnosis, length of stay, downstream test volume. A model can be highly accurate and change nothing, because the constraint was somewhere else in the process. It can also be accurate, change behaviour, and make things worse through alert fatigue.

Regulation, at a level of generality I'm comfortable defending: software that informs clinical decisions is treated as a medical device in most major jurisdictions, with evidence requirements scaled to the risk of the decision it touches. The load-bearing document is the intended-use statement, because it defines what the product is permitted to claim. A tool authorised to prioritise a worklist is not authorised to tell anyone whether they have a disease, and building a UI that implies otherwise creates a real problem no matter how good the AUC is. Two consequences follow for engineering teams. First, retraining is a change to the device, so a plan to ship model updates monthly needs a change-control story before the first release, not after. Second, post-market monitoring is part of the product: track input distributions, score distributions, alarm rate and override rate, because a drifting alarm rate is often the earliest available signal of dataset shift and it costs you nothing in labels to watch. I'm an engineer and not a regulatory specialist, jurisdictions differ, and the useful move is to involve people who do this professionally while the architecture is still soft.

Questions that separate a real project from a demo

If you're evaluating a clinical AI proposal — your own or a vendor's — these six questions do more work than any amount of architecture discussion, and an unconvincing answer to any of them is worth stopping over.

  1. What is the prevalence in the population that will actually be tested, and what PPV does your chosen operating point imply at that prevalence? No answer here means there is no product specification yet.
  2. Who owns the threshold, and can a site change it after deployment without retraining?
  3. What happens to a flagged case, and what happens to a cleared one? Write both paths end to end, naming who is accountable at each step.
  4. Has the model been evaluated on data from a site it never trained on, preferably with different equipment? If not, generalisation is an assumption.
  5. What are the per-subgroup numbers with intervals, and how many positive cases sit in the smallest subgroup?
  6. What signal would tell you it has stopped working, and would anyone notice within a week?

The failure that keeps repeating is not a bad model. It's a good model shipped against an unstated population, at an operating point nobody chose deliberately, validated on the site it was born in, and monitored by nothing but complaints. All four of those are evaluation problems, and all four are fixable with arithmetic and discipline rather than a larger network.

None of this is guidance for interpreting anyone's actual test result — that conversation belongs with a clinician who knows the patient. But the constraints above are not a reason to be dismissive about medical AI. They are the specification. The teams making real progress are the ones who treated prevalence, thresholds, external validation and subgroup reporting as the core of the work rather than as paperwork bolted on at the end.