Course Content
Evaluating and Testing GenAI Models
4 sections · 13 lessons
Quantitative vs. Qualitative Hallucination Measures
Two models are evaluated on the same 500 claims from a medical information assistant. The report says:
Model A hallucination rate 8.0% (40 / 500)Model B hallucination rate 6.0% (30 / 500)B wins by two points, so B ships. Then someone opens the annotation file and looks at what the errors were.
| Severity | Definition | Model A | Model B |
|---|---|---|---|
| Critical | Could cause direct harm if acted on — wrong dosage, invented contraindication | 3 | 8 |
| Major | Materially misleading — wrong mechanism, wrong study finding | 11 | 9 |
| Minor | Imprecise but not misleading — approximate date, rounded figure | 26 | 13 |
| Total | 40 | 30 |
Model B produces fewer hallucinations and nearly three times as many critical ones. Its lower rate came almost entirely from cleaning up minor imprecision. On a medical assistant, that is the opposite of an improvement, and the headline number said "ship it".
Quantitative and qualitative measurement are not two styles or two preferences. They answer different questions — how much versus what kind and how bad — and a system evaluated on only one of them has a specific, predictable blind spot.
The quantitative side: what the numbers are and what each one hides
| Metric | Definition | Answers | Hides |
|---|---|---|---|
| Hallucination rate | Unsupported claims / total claims | How often does it happen? | Severity; whether errors cluster in a few responses |
| Faithfulness | 1 − hallucination rate, against a given source | How grounded is the output? | Whether omissions matter; open-world truth |
| Contradiction rate | Claims that contradict the source / total | How often does it assert the opposite? | Nothing much — this is the cleanest of the set |
| Factual accuracy | Correct claims / verifiable claims, against the world | Is it right? | Retrieval coverage; unverifiable claims silently excluded |
| Clean-response rate | Responses with zero hallucinated claims / responses | What does a user actually experience? | How bad the flawed responses were |
| Self-consistency | Agreement across k resamples of the same prompt | Is the model uncertain? | Stable false beliefs, which score perfectly |
| Calibration error | Gap between stated confidence and observed accuracy | Can we trust its hedging? | Nothing — but almost nobody measures it |
Two of these deserve full worked treatment because they are the ones most often computed wrongly.
Clean-response rate versus claim rate
These diverge in a way that tells you about the distribution of errors. Take 120 responses containing 480 claims, with 38 hallucinations.
If those 38 hallucinations are spread across 31 different responses:
If instead they were concentrated in 11 responses:
Same 7.9% claim rate; a 16-point difference in the fraction of users who receive a fully reliable answer. The concentrated case is much better for a product and points at a specific trigger — a query type, a document format, a topic — that you can go and find. The spread case means the model is unreliable everywhere at a low rate, which is a harder and more fundamental problem. Always report both, and treat their ratio as a diagnostic.
Calibration: the metric almost nobody computes
If a model hedges ("I believe", "approximately", "this may be"), those hedges are only useful if they track reality. Expected calibration error measures whether they do. Bin the claims by the model's stated or estimated confidence, and compare average confidence to observed accuracy within each bin.
Worked example over 500 claims in five bins:
| Confidence bin | nb | Mean confidence | Observed accuracy | Gap | nb× gap |
|---|---|---|---|---|---|
| 0.50 – 0.60 | 80 | 0.55 | 0.50 | 0.05 | 4.0 |
| 0.60 – 0.70 | 110 | 0.65 | 0.58 | 0.07 | 7.7 |
| 0.70 – 0.80 | 130 | 0.75 | 0.66 | 0.09 | 11.7 |
| 0.80 – 0.90 | 110 | 0.85 | 0.74 | 0.11 | 12.1 |
| 0.90 – 1.00 | 70 | 0.95 | 0.81 | 0.14 | 9.8 |
| Total | 500 | 45.3 |
An ECE of 9.1 points is poor but the shape is the real finding: the gap grows monotonically with confidence, from 0.05 to 0.14. The model is most overconfident exactly where a user is most likely to skip verification. A system that routes "high confidence" answers past human review is therefore routing its worst-calibrated outputs around its only safety net.
An uncalibrated confidence score is worse than no confidence score, because it will be used to decide what not to check.
Severity weighting turns counts into consequences
Return to the opening. Assign weights by consequence — 10 for critical, 4 for major, 1 for minor — and compute a harm score:
Model A's harm score is 100 against B's 129 — a 29% advantage for the model with the worse raw rate. The ranking flips, and it flips because the weights encode something the raw count refuses to: that the errors are not interchangeable.
The weights are a judgement call and should be argued about openly, by the people who own the consequences, and written down. That argument is far more valuable than the number it produces. A team that has agreed "an invented dosage is worth ten rounded dates" has done real work on what their product is for.
The qualitative side: structure, not vibes
"Qualitative" does not mean unstructured impressions. It means capturing properties of individual failures that a count cannot express, in a form consistent enough to aggregate later. A qualitative record has fields.
1{2 "id": "h-2026-0412",3 "response_id": "resp_88f21",4 "claim": "Ibuprofen is safe to combine with warfarin at standard doses.",5 "verdict": "CONTRADICTED",6 "type": "fabricated_safety_claim",7 "severity": "critical",8 "detectable_by_user": false,9 "trigger_hypothesis": "query contained two drug names and no explicit interaction question",10 "source_evidence": "formulary-2026, section 4.2 (documents a major interaction)",11 "would_have_been_caught_by": ["drug-interaction-lookup", "expert-review"],12 "proposed_check": "flag any response naming 2+ drugs without an interaction lookup",13 "annotator": "clinical-reviewer-3",14 "date": "2026-04-12"15}The four fields that earn their place:
detectable_by_user— this, more than severity, determines the design response. An error the user spots immediately is self-correcting; an invisible one is not.trigger_hypothesis— a guess at what caused it. Aggregate these and patterns appear that no metric surfaces: "seven of our nine critical errors involved a query with two entities and an implicit comparison."would_have_been_caught_by— turns the annotation into a specification for the next automated check.proposed_check— the concrete artefact. This is the field that makes qualitative work compound rather than accumulate in a document nobody reads.
What only qualitative analysis can find
| Finding | Why no metric shows it |
|---|---|
| A new failure mode after a prompt change | Metrics only count categories you defined in advance |
| Errors concentrated on one customer segment | Aggregate rates average across segments |
| The same wrong fact recurring from one bad source document | Counted as many independent errors, not one root cause |
| Errors that are technically correct but useless | They pass every faithfulness check |
| Systematic omission of a required caveat | Nothing false was said; the metric is silent on what is missing |
| The evaluation rubric itself being wrong | A metric cannot audit its own definition |
That last row is the one people underestimate. Metrics measure what they measure, forever, with no mechanism for noticing they have become invalid. The only thing that catches a broken metric is a person reading outputs and saying "this scored 0.94 and it is bad."
Putting them together
| Quantitative | Qualitative | |
|---|---|---|
| Answers | How much, trending which way | What kind, how bad, why |
| Scales to | Millions of claims | Tens to low hundreds |
| Reproducible | Yes | Only with a strong codebook |
| Finds new failure modes | Never | This is its purpose |
| Supports a ship/no-ship gate | Yes, with intervals | Yes, as a veto on severity |
| Cost per item | Fractions of a cent | Minutes of expert time |
| Degrades by | Measuring the wrong construct, silently | Coder drift and small samples |
| Run it | Continuously | Weekly on a sample, plus after every incident |
The productive arrangement is a loop with a specific direction of travel:
Quantitative sweep over all traffic | | select: worst-scoring 3%, plus a RANDOM 1% audit vQualitative coding of that sample | | produces: type, severity, trigger, proposed check vNew automated checks + severity weights | | feed back into vQuantitative sweep (now measuring one more thing than last month)The random 1% is doing the important work, and it is what teams drop first when budgets tighten. If humans only review what the metrics already flagged, the qualitative layer can only ever confirm what you already knew. The random stream is the only channel through which an unknown failure mode can enter the system.
Reviewing only flagged outputs makes your evaluation self-confirming. The random sample is what keeps it capable of surprising you.
Tracking hallucination over time
A rate measured once is nearly useless; the operational question is whether it is moving. But daily rates are noisy, and naive thresholds are badly matched to the kind of change that actually happens.
Baseline hallucination rate 5.0%, measured daily on 400 claims. The daily standard deviation is:
A conventional 3-sigma daily alarm therefore fires at 0.05+3(0.0109)=0.0827, i.e. 8.27%. Now watch a week of gradual degradation after a prompt change:
| Day | Observed rate | EWMA (λ=0.3) | EWMA limit | Daily 3σ alarm? |
|---|---|---|---|---|
| 1 | 4.8% | 4.94% | 6.37% | No |
| 2 | 5.2% | 5.02% | 6.37% | No |
| 3 | 6.1% | 5.34% | 6.37% | No |
| 4 | 5.8% | 5.48% | 6.37% | No |
| 5 | 7.0% | 5.94% | 6.37% | No |
| 6 | 7.5% | 6.41% | 6.37% | No |
| 7 | 8.2% | 6.94% | 6.37% | No |
The exponentially weighted moving average is Et=λxt+(1−λ)Et−1, and its control limit is:
The EWMA crosses its limit on day 6. The daily 3-sigma rule never fires at all, not even on day 7 at 8.2%, because a single day at 8.2% is genuinely within the noise of a 5% process at n=400. Both rules are behaving correctly; they are just tuned for different failures. Sudden breakage needs the daily rule; gradual drift — which is what prompt changes, corpus drift and model updates actually produce — needs the EWMA. Run both.
1import numpy as np23def ewma_monitor(daily_rates, p0, n_per_day, lam=0.3, L=3.0):4 sigma = np.sqrt(p0 * (1 - p0) / n_per_day)5 ucl = p0 + L * sigma * np.sqrt(lam / (2 - lam))6 lcl = p0 - L * sigma * np.sqrt(lam / (2 - lam))7 e, out = p0, []8 for day, x in enumerate(daily_rates, start=1):9 e = lam * x + (1 - lam) * e10 out.append({"day": day, "rate": x, "ewma": round(e, 5),11 "alarm": bool(e > ucl or e < lcl)})12 return {"ucl": round(ucl, 5), "lcl": round(lcl, 5), "series": out}A falling EWMA deserves the same scrutiny as a rising one. A hallucination rate that drops sharply is occasionally a genuine improvement and is more often a broken detector, a changed traffic mix, or a model that has started refusing more questions. Alarm on both directions.
What this means when you own a generation system
Put severity in the schema on day one. Retro-fitting severity onto months of binary hallucinated/not-hallucinated annotations means re-annotating everything, and until you do, every comparison you make has the flaw from the opening: it can rank a model with three times the critical-error count as the winner. A three-level severity field costs the annotator four seconds per item and is the difference between a number you can act on and a number that can mislead you.
Budget qualitative review as a fixed recurring cost, not as something you do when a metric goes red. Two hours a week of an expert reading a random sample of production outputs will find things your entire automated suite cannot, and each finding converts into a permanent check. Teams that only do qualitative work reactively, during incidents, get the same information months later and at much higher cost.
And publish the two views together, always. A dashboard that shows "hallucination rate: 6.0%, down from 8.0%" and nothing else is technically accurate and, in the opening scenario, actively harmful. The version that shows the rate and the severity breakdown and the harm score and the interval on each is the version where the person reading it can see that something has gone badly wrong in the only category that matters.