Evaluating and Testing GenAI Models

Quantitative vs. Qualitative Hallucination Measures


Two models are evaluated on the same 500 claims from a medical information assistant. The report says:

Text
Model A   hallucination rate  8.0%   (40 / 500)Model B   hallucination rate  6.0%   (30 / 500)

B wins by two points, so B ships. Then someone opens the annotation file and looks at what the errors were.

SeverityDefinitionModel AModel B
CriticalCould cause direct harm if acted on — wrong dosage, invented contraindication38
MajorMaterially misleading — wrong mechanism, wrong study finding119
MinorImprecise but not misleading — approximate date, rounded figure2613
Total4030

Model B produces fewer hallucinations and nearly three times as many critical ones. Its lower rate came almost entirely from cleaning up minor imprecision. On a medical assistant, that is the opposite of an improvement, and the headline number said "ship it".

Quantitative and qualitative measurement are not two styles or two preferences. They answer different questions — how much versus what kind and how bad — and a system evaluated on only one of them has a specific, predictable blind spot.

Two models, 500 claims, and the wrong winner8.0 percent3261006.0 percent813129Hallucination rateCritical errorsMinor errorsHarm scoreModel AModel BWeighting critical 10, major 4 and minor 1 turns the counts into a harm score the headline rate erases.
B leads on the headline rate and is the more dangerous model, because rates count errors while consequences weigh them.

The quantitative side: what the numbers are and what each one hides

MetricDefinitionAnswersHides
Hallucination rateUnsupported claims / total claimsHow often does it happen?Severity; whether errors cluster in a few responses
Faithfulness1 − hallucination rate, against a given sourceHow grounded is the output?Whether omissions matter; open-world truth
Contradiction rateClaims that contradict the source / totalHow often does it assert the opposite?Nothing much — this is the cleanest of the set
Factual accuracyCorrect claims / verifiable claims, against the worldIs it right?Retrieval coverage; unverifiable claims silently excluded
Clean-response rateResponses with zero hallucinated claims / responsesWhat does a user actually experience?How bad the flawed responses were
Self-consistencyAgreement across kk resamples of the same promptIs the model uncertain?Stable false beliefs, which score perfectly
Calibration errorGap between stated confidence and observed accuracyCan we trust its hedging?Nothing — but almost nobody measures it

Two of these deserve full worked treatment because they are the ones most often computed wrongly.

Clean-response rate versus claim rate

These diverge in a way that tells you about the distribution of errors. Take 120 responses containing 480 claims, with 38 hallucinations.

Claim-level rate=38480=7.9%\text{Claim-level rate} = \frac{38}{480} = 7.9\%

If those 38 hallucinations are spread across 31 different responses:

Clean-response rate=120−31120=89120=74.2%\text{Clean-response rate} = \frac{120 - 31}{120} = \frac{89}{120} = 74.2\%

If instead they were concentrated in 11 responses:

Clean-response rate=120−11120=109120=90.8%\text{Clean-response rate} = \frac{120 - 11}{120} = \frac{109}{120} = 90.8\%

Same 7.9% claim rate; a 16-point difference in the fraction of users who receive a fully reliable answer. The concentrated case is much better for a product and points at a specific trigger — a query type, a document format, a topic — that you can go and find. The spread case means the model is unreliable everywhere at a low rate, which is a harder and more fundamental problem. Always report both, and treat their ratio as a diagnostic.

Calibration: the metric almost nobody computes

If a model hedges ("I believe", "approximately", "this may be"), those hedges are only useful if they track reality. Expected calibration error measures whether they do. Bin the claims by the model's stated or estimated confidence, and compare average confidence to observed accuracy within each bin.

ECE=∑b=1BnbN∣acc(b)−conf(b)∣\text{ECE} = \sum_{b=1}^{B} \frac{n_b}{N}\left|\text{acc}(b) - \text{conf}(b)\right|

Worked example over 500 claims in five bins:

Confidence binnbn_bMean confidenceObserved accuracyGapnb×n_b \times gap
0.50 – 0.60800.550.500.054.0
0.60 – 0.701100.650.580.077.7
0.70 – 0.801300.750.660.0911.7
0.80 – 0.901100.850.740.1112.1
0.90 – 1.00700.950.810.149.8
Total50045.3
ECE=45.3500=0.091\text{ECE} = \frac{45.3}{500} = 0.091

An ECE of 9.1 points is poor but the shape is the real finding: the gap grows monotonically with confidence, from 0.05 to 0.14. The model is most overconfident exactly where a user is most likely to skip verification. A system that routes "high confidence" answers past human review is therefore routing its worst-calibrated outputs around its only safety net.

An uncalibrated confidence score is worse than no confidence score, because it will be used to decide what not to check.

Severity weighting turns counts into consequences

Return to the opening. Assign weights by consequence — 10 for critical, 4 for major, 1 for minor — and compute a harm score:

H=∑sws⋅nsH = \sum_{s} w_s \cdot n_s

HA=(10×3)+(4×11)+(1×26)=30+44+26=100H_A = (10 \times 3) + (4 \times 11) + (1 \times 26) = 30 + 44 + 26 = 100
HB=(10×8)+(4×9)+(1×13)=80+36+13=129H_B = (10 \times 8) + (4 \times 9) + (1 \times 13) = 80 + 36 + 13 = 129

Model A's harm score is 100 against B's 129 — a 29% advantage for the model with the worse raw rate. The ranking flips, and it flips because the weights encode something the raw count refuses to: that the errors are not interchangeable.

The weights are a judgement call and should be argued about openly, by the people who own the consequences, and written down. That argument is far more valuable than the number it produces. A team that has agreed "an invented dosage is worth ten rounded dates" has done real work on what their product is for.

The qualitative side: structure, not vibes

"Qualitative" does not mean unstructured impressions. It means capturing properties of individual failures that a count cannot express, in a form consistent enough to aggregate later. A qualitative record has fields.

JSON
{  "id": "h-2026-0412",  "response_id": "resp_88f21",  "claim": "Ibuprofen is safe to combine with warfarin at standard doses.",  "verdict": "CONTRADICTED",  "type": "fabricated_safety_claim",  "severity": "critical",  "detectable_by_user": false,  "trigger_hypothesis": "query contained two drug names and no explicit interaction question",  "source_evidence": "formulary-2026, section 4.2 (documents a major interaction)",  "would_have_been_caught_by": ["drug-interaction-lookup", "expert-review"],  "proposed_check": "flag any response naming 2+ drugs without an interaction lookup",  "annotator": "clinical-reviewer-3",  "date": "2026-04-12"}

The four fields that earn their place:

  • detectable_by_user — this, more than severity, determines the design response. An error the user spots immediately is self-correcting; an invisible one is not.
  • trigger_hypothesis — a guess at what caused it. Aggregate these and patterns appear that no metric surfaces: "seven of our nine critical errors involved a query with two entities and an implicit comparison."
  • would_have_been_caught_by — turns the annotation into a specification for the next automated check.
  • proposed_check — the concrete artefact. This is the field that makes qualitative work compound rather than accumulate in a document nobody reads.

What only qualitative analysis can find

FindingWhy no metric shows it
A new failure mode after a prompt changeMetrics only count categories you defined in advance
Errors concentrated on one customer segmentAggregate rates average across segments
The same wrong fact recurring from one bad source documentCounted as many independent errors, not one root cause
Errors that are technically correct but uselessThey pass every faithfulness check
Systematic omission of a required caveatNothing false was said; the metric is silent on what is missing
The evaluation rubric itself being wrongA metric cannot audit its own definition

That last row is the one people underestimate. Metrics measure what they measure, forever, with no mechanism for noticing they have become invalid. The only thing that catches a broken metric is a person reading outputs and saying "this scored 0.94 and it is bad."

Putting them together

QuantitativeQualitative
AnswersHow much, trending which wayWhat kind, how bad, why
Scales toMillions of claimsTens to low hundreds
ReproducibleYesOnly with a strong codebook
Finds new failure modesNeverThis is its purpose
Supports a ship/no-ship gateYes, with intervalsYes, as a veto on severity
Cost per itemFractions of a centMinutes of expert time
Degrades byMeasuring the wrong construct, silentlyCoder drift and small samples
Run itContinuouslyWeekly on a sample, plus after every incident

The productive arrangement is a loop with a specific direction of travel:

Text
Quantitative sweep over all traffic        |        |  select: worst-scoring 3%, plus a RANDOM 1% audit        vQualitative coding of that sample        |        |  produces: type, severity, trigger, proposed check        vNew automated checks + severity weights        |        |  feed back into        vQuantitative sweep  (now measuring one more thing than last month)

The random 1% is doing the important work, and it is what teams drop first when budgets tighten. If humans only review what the metrics already flagged, the qualitative layer can only ever confirm what you already knew. The random stream is the only channel through which an unknown failure mode can enter the system.

Reviewing only flagged outputs makes your evaluation self-confirming. The random sample is what keeps it capable of surprising you.

Tracking hallucination over time

A rate measured once is nearly useless; the operational question is whether it is moving. But daily rates are noisy, and naive thresholds are badly matched to the kind of change that actually happens.

Baseline hallucination rate 5.0%, measured daily on 400 claims. The daily standard deviation is:

σ=0.05×0.95400=0.00011875=0.0109\sigma = \sqrt{\frac{0.05 \times 0.95}{400}} = \sqrt{0.00011875} = 0.0109

A conventional 3-sigma daily alarm therefore fires at 0.05+3(0.0109)=0.08270.05 + 3(0.0109) = 0.0827, i.e. 8.27%. Now watch a week of gradual degradation after a prompt change:

DayObserved rateEWMA (λ=0.3\lambda = 0.3)EWMA limitDaily 3σ alarm?
14.8%4.94%6.37%No
25.2%5.02%6.37%No
36.1%5.34%6.37%No
45.8%5.48%6.37%No
57.0%5.94%6.37%No
67.5%6.41%6.37%No
78.2%6.94%6.37%No

The exponentially weighted moving average is Et=λxt+(1−λ)Et−1E_t = \lambda x_t + (1-\lambda)E_{t-1}, and its control limit is:

UCL=p0+3σλ2−λ=0.05+3(0.0109)0.31.7=0.05+0.0327×0.420=0.0637\text{UCL} = p_0 + 3\sigma\sqrt{\frac{\lambda}{2 - \lambda}} = 0.05 + 3(0.0109)\sqrt{\frac{0.3}{1.7}} = 0.05 + 0.0327 \times 0.420 = 0.0637

The EWMA crosses its limit on day 6. The daily 3-sigma rule never fires at all, not even on day 7 at 8.2%, because a single day at 8.2% is genuinely within the noise of a 5% process at n=400n = 400. Both rules are behaving correctly; they are just tuned for different failures. Sudden breakage needs the daily rule; gradual drift — which is what prompt changes, corpus drift and model updates actually produce — needs the EWMA. Run both.

Python
import numpy as npdef ewma_monitor(daily_rates, p0, n_per_day, lam=0.3, L=3.0):    sigma = np.sqrt(p0 * (1 - p0) / n_per_day)    ucl = p0 + L * sigma * np.sqrt(lam / (2 - lam))    lcl = p0 - L * sigma * np.sqrt(lam / (2 - lam))    e, out = p0, []    for day, x in enumerate(daily_rates, start=1):        e = lam * x + (1 - lam) * e        out.append({"day": day, "rate": x, "ewma": round(e, 5),                    "alarm": bool(e > ucl or e < lcl)})    return {"ucl": round(ucl, 5), "lcl": round(lcl, 5), "series": out}

A falling EWMA deserves the same scrutiny as a rising one. A hallucination rate that drops sharply is occasionally a genuine improvement and is more often a broken detector, a changed traffic mix, or a model that has started refusing more questions. Alarm on both directions.

What this means when you own a generation system

Put severity in the schema on day one. Retro-fitting severity onto months of binary hallucinated/not-hallucinated annotations means re-annotating everything, and until you do, every comparison you make has the flaw from the opening: it can rank a model with three times the critical-error count as the winner. A three-level severity field costs the annotator four seconds per item and is the difference between a number you can act on and a number that can mislead you.

Budget qualitative review as a fixed recurring cost, not as something you do when a metric goes red. Two hours a week of an expert reading a random sample of production outputs will find things your entire automated suite cannot, and each finding converts into a permanent check. Teams that only do qualitative work reactively, during incidents, get the same information months later and at much higher cost.

And publish the two views together, always. A dashboard that shows "hallucination rate: 6.0%, down from 8.0%" and nothing else is technically accurate and, in the opening scenario, actively harmful. The version that shows the rate and the severity breakdown and the harm score and the interval on each is the version where the person reading it can see that something has gone badly wrong in the only category that matters.