Evaluating and Testing GenAI Models

Common Hallucination Types (Fabricated Facts, Reasoning Errors)


In 2023 a federal court in New York sanctioned two lawyers who had filed a brief citing six judicial decisions that did not exist. The citations were not garbled or misremembered. They had case names, reporter volumes, page numbers, courts, years, and quoted passages — all in the correct format, all fabricated. When the opposing counsel could not find them, one of the lawyers asked the model whether the cases were real. It said yes.

What makes this the canonical example is not the embarrassment; it is the shape of the failure. The output was not degraded, garbled, or obviously uncertain. It was maximally fluent and maximally confident precisely where it was maximally wrong. That inversion — confidence uncorrelated with correctness — is the defining property of hallucination and the reason it cannot be caught by the intuitions that catch ordinary software bugs.

To measure hallucination you first have to be able to name what kind you are looking at, because the detection method, the metric, and the fix are different for each kind. A single "hallucination rate" that pools all of them together tells you nothing actionable.

The taxonomy that decides how you detect itUnsupported outputIntrinsic —contradicts the sourceExtrinsic — nosource says itEntity swappedNumber inventedCitation fabricatedConclusion unearned
The six fake cases were formatted perfectly — fluency is what makes extrinsic fabrication survive review.

What a hallucination is, precisely

A hallucination is content the model presents as fact that is either unsupported by the provided source or false about the world, generated with no signal distinguishing it from content that is supported or true.

Three parts of that definition earn their place:

  • Presented as fact. A model writing fiction on request is not hallucinating. The failure requires an implicit assertion of truth.
  • Unsupported or false. These are different failures, and conflating them is the most common measurement error in the field. More on this below.
  • No distinguishing signal. This is what makes it hard. The model does not encode "I am now making this up". Token probabilities on a fabricated citation often look much like those on a real one.

Why it happens at all

Not mysticism — four concrete mechanisms:

MechanismWhat it doesSymptom it produces
Next-token training objectiveThe model is trained to produce plausible continuations, never verified ones. Nothing in the loss distinguishes true from true-sounding.Fluent fabrication in the exact style of correct answers
Parametric compressionFacts are stored as distributed weights, not as retrievable records. Rare facts are stored approximately, and approximate storage of a name or a date reconstructs as a different name or date, not as a blank.Entity confusion; plausible wrong numbers
Pattern completion pressureHaving emitted "According to a study published in", the model is now under enormous distributional pressure to emit a journal name. There is no token for "actually I do not know one".Fabricated citations, invented statistics
Alignment for helpfulnessPreference training rewards answers users rated highly. Users rate confident, complete answers above hedged or refusing ones.Overconfidence; reluctance to say "I do not know"

A model has no internal representation of the difference between remembering and inventing. Both are the same operation — sampling a plausible continuation — and only one of them happens to be right.

The taxonomy that matters: intrinsic versus extrinsic

The primary split is defined relative to a source document, if the task provides one.

Intrinsic hallucinationExtrinsic hallucination
DefinitionThe output contradicts the sourceThe output makes a claim the source does not address
ExampleSource says revenue rose 12%; summary says revenue rose 21%Source says revenue rose 12%; summary adds "driven by strong European demand" (not in source)
Detectable byEntailment against the source alone — fully decidableEntailment gives "neutral"; you need external verification to know if it is true
Always an error?Yes, unconditionallyNo — it may be true and useful, or true and out of scope, or false
Typical fixBetter grounding, constrained decoding, copy mechanismsExplicit scope instructions; citation requirements

This distinction is not academic. Intrinsic hallucinations can be detected automatically at scale with an NLI model and no internet access. Extrinsic ones cannot — they require retrieval, and retrieval failure is indistinguishable from falsehood unless you handle it deliberately. Pooling them into one rate makes the metric un-actionable, because half of it is cheap to fix and half is not.

When there is no source document — open-ended question answering — the analogous split is closed-domain (the model contradicts something in its context or instructions) versus open-domain (the model asserts something false about the world).

The failure types you will actually see

Fabricated entities and citations

The model invents a source, paper, product, API, person, or legal case with correct-looking metadata. This is the most dangerous type because the format is the credibility signal, and the format is perfect.

Text
Prompt:  "Cite a study on the effect of sleep on memory consolidation."Output:  "Ratcliffe & Moreno (2019), 'Sleep-Dependent Consolidation of          Declarative Memory', Journal of Cognitive Neuroscience 31(4),          pp. 612-628, found a 34% improvement in recall."Reality: The journal exists. The volume/issue numbering is plausible.         The authors, title, page range and finding do not exist.

Detection is unusually easy for this one if you build for it: resolve every citation against a real index (Crossref, PubMed, a package registry, a case database). A citation that does not resolve is a hard failure with no judgement call required. Any system that emits references and does not do this is choosing not to catch its most catastrophic error class.

Named entity substitution

The model produces the right kind of thing with the wrong identity: the correct role attributed to a different person, a real quotation attributed to the wrong speaker, the right event in the wrong city. This comes directly from the compression mechanism — nearby entities in embedding space are interchangeable under approximate recall.

Numeric and statistical fabrication

Numbers are the highest-risk tokens in any output. They are short, they carry disproportionate meaning, they cannot be sanity-checked by reading fluency, and models produce them with total confidence. "Roughly 40% of respondents" appearing in a summary of a source that reported 27% is invisible to every fluency-based check and fatal to a decision made on it.

Reasoning errors dressed as conclusions

Distinct from factual fabrication: every premise is true, and the inference is invalid.

Text
Premises (all true):   All members of the committee are economists.                       Dr Okafor is an economist.Stated conclusion:     Therefore Dr Okafor is on the committee.

Common sub-forms: affirming the consequent (as above), unit errors in arithmetic chains, off-by-one in date arithmetic, and confusing correlation with causation while citing a correct study. Fact-checking each claim individually passes all of them, because each individual claim is true. Only checking the inference catches it — which is why claim-level verification alone gives a falsely reassuring picture on reasoning-heavy tasks.

False attribution

Real content, wrong origin: a genuine finding credited to the wrong study, a real quote to the wrong author, a real API method to the wrong library. Especially insidious in technical documentation, where pandas.DataFrame.explode_all() reads exactly like something that ought to exist.

Anachronism and temporal error

Claims that are true of one time asserted about another: the model states current officeholders, prices, versions, or population figures from its training data as though they hold now, or places an event before something it depended on. Any deployment answering time-sensitive questions needs an explicit temporal check, not just a factual one.

Source unfaithfulness in grounded tasks

Given a document and asked to summarise, the model adds information from its parametric memory. This is often true information — which is what makes it hard to argue about — but it is out of scope, unverifiable by the reader against the document, and destroys the guarantee that makes grounded generation useful in the first place.

Partial truth

The hardest type to score. The claim is 80% right with a wrong detail embedded: correct mechanism, wrong dosage; correct API, wrong argument order; correct historical event, wrong year. Binary claim-level labelling forces raters into an arbitrary call, and the resulting inter-rater agreement is poor. Handle it by decomposing further — "correct mechanism" and "correct dosage" are two claims, and scoring them separately eliminates the ambiguity.

What it costs you, by domain

DomainTypical hallucinationCost of one instanceDetectability
Software documentationNon-existent method or flagLow — fails immediately, developer investigatesHigh: execute it
LegalFabricated case citationSevere — sanctions, malpracticeHigh if you resolve citations; otherwise nil
MedicalWrong dosage, invented contraindicationCatastrophic and irreversibleLow — reads exactly like correct guidance
FinancialInvented figure in a summary of a filingSevere — decisions and disclosure riskMedium: numbers are checkable against the source
Customer supportInvented policy or entitlementModerate — but you may be held to itMedium: check against the policy corpus
Creative writingInvented detailNone — it is the taskNot applicable

Note the pattern in column four: the domains where hallucination is most costly are the ones where it is least detectable by a non-expert reader. That inverse relationship is the whole argument for building verification into the system rather than relying on the user to notice.

Quantifying hallucination correctly

The base metric is claim-level. Decompose each response into atomic claims, label each, and count.

Hallucination rate=unsupported or false claimstotal claimsFaithfulness=1−hallucination rate\text{Hallucination rate} = \frac{\text{unsupported or false claims}}{\text{total claims}} \qquad \text{Faithfulness} = 1 - \text{hallucination rate}

Worked example: 50 responses yield 200 atomic claims. 23 are unsupported.

p^=23200=0.115faithfulness=0.885\hat{p} = \frac{23}{200} = 0.115 \qquad \text{faithfulness} = 0.885

The clustering correction that almost everyone omits

Those 200 claims are not 200 independent observations. They come from 50 responses, about 4 claims each, and claims within a response are correlated — a response that goes off the rails tends to go off the rails several times. Treating them as independent understates your uncertainty.

The naive standard error:

SEnaive=0.115×0.885200=0.000509=0.0226    (2.26 pp)SE_{\text{naive}} = \sqrt{\frac{0.115 \times 0.885}{200}} = \sqrt{0.000509} = 0.0226 \;\;(2.26 \text{ pp})

The design effect for clustered sampling, with average cluster size m=4m = 4 and intra-cluster correlation ρ=0.3\rho = 0.3:

DEFF=1+(m−1)ρ=1+3×0.3=1.9\text{DEFF} = 1 + (m - 1)\rho = 1 + 3 \times 0.3 = 1.9
neff=2001.9=105SEtrue=0.115×0.885105=0.000969=0.0311    (3.11 pp)n_{\text{eff}} = \frac{200}{1.9} = 105 \qquad SE_{\text{true}} = \sqrt{\frac{0.115 \times 0.885}{105}} = \sqrt{0.000969} = 0.0311 \;\;(3.11 \text{ pp})

Your 95% interval on the hallucination rate is therefore 11.5%±6.1%11.5\% \pm 6.1\%, or roughly [5.4%, 17.6%][5.4\%,\ 17.6\%] — not the [7.1%, 15.9%][7.1\%,\ 15.9\%] the naive calculation suggests. A competing model at 8.9% is comfortably inside that interval and has not been shown to be better.

Claims are clustered inside responses. Ignore that and you will report differences as real that are entirely within the noise of which fifty responses you happened to sample.

The practical shortcut, if estimating ρ\rho feels like overreach: bootstrap by response, not by claim. Resample whole responses with replacement, recompute the rate, and read the percentiles. It handles the clustering without requiring you to estimate anything.

Python
import numpy as npdef hallucination_rate_ci(responses, iters=2000, seed=0):    """responses: list of lists of 0/1 labels, one list per response.       1 = hallucinated claim. Bootstraps at the RESPONSE level."""    rng = np.random.default_rng(seed)    flat = [c for r in responses for c in r]    point = sum(flat) / len(flat)    n = len(responses)    draws = []    for _ in range(iters):        idx = rng.integers(0, n, n)        sampled = [c for i in idx for c in responses[i]]        draws.append(sum(sampled) / len(sampled))    lo, hi = np.percentile(draws, [2.5, 97.5])    return round(point, 4), round(lo, 4), round(hi, 4)

Report it broken down, never pooled

A single 11.5% figure hides everything you need to act on. The useful table:

TypeCountRateSeverityFixable by
Intrinsic (contradicts source)73.5%HighGrounding, decoding constraints
Extrinsic, verified false42.0%HighRetrieval + refusal on low evidence
Extrinsic, unverifiable94.5%MediumScope instructions, citation requirement
Reasoning error31.5%HighDecomposition, verification of steps

Now the engineering priority is obvious: intrinsic contradictions are the largest high-severity bucket and the cheapest to detect, so they go first. The pooled 11.5% would not have told you that.

Detecting them: red flags and real methods

Fast heuristics for a human reviewer or a cheap pre-filter:

  • Any specific number, date, percentage, or measurement — the highest-risk token class
  • Any proper noun that could be looked up: person, paper, case, product, API symbol
  • Superlatives and absolutes: "the first", "the only", "always", "never"
  • Confident answers to questions containing a false presupposition ("Why does aspirin cure tinnitus?")
  • Suspiciously round statistics: "approximately 40% of users" with no source
  • Answers to questions about events after the model's training cutoff

Real methods, with their trade-offs:

MethodHow it worksCatchesCostLimitation
NLI against sourceEntailment model on (source, claim)Intrinsic onlyVery lowSilent on anything the source omits
Citation resolutionLook every reference up in a real indexFabricated sourcesVery lowOnly applies to referenced claims
Retrieval + verificationSearch, then judge the claim against retrieved evidenceOpen-domain falsehoodMediumRetrieval misses read as falsehood
Self-consistency samplingSample the same prompt kk times; measure agreement across samplesLow-confidence fabricationk×k\times generation costConsistently-held false beliefs pass cleanly
Token-probability signalsFlag spans with low mean token probability or high entropySome uncertaintyNear zeroWeak correlation; needs logit access
Execution / toolingRun the code, call the API, evaluate the arithmeticEverything decidableLowOnly for executable claims

Self-consistency deserves a caution. It measures whether the model is stable, not whether it is right. If the model has a firmly encoded false belief — a common misconception that appeared thousands of times in training — it will produce the same wrong answer at every temperature, and self-consistency will pronounce it reliable. Stability and truth are different properties, and this method only sees the first.

Evaluating your detector, not just your model

A detector needs its own numbers. Suppose across 500 claims, 60 are genuinely hallucinated. Your detector flags 80, of which 45 are true hallucinations.

Precision=4580=0.5625Recall=4560=0.750\text{Precision} = \frac{45}{80} = 0.5625 \qquad \text{Recall} = \frac{45}{60} = 0.750

F1=2×0.5625×0.7500.5625+0.750=0.843751.3125=0.643F_1 = \frac{2 \times 0.5625 \times 0.750}{0.5625 + 0.750} = \frac{0.84375}{1.3125} = 0.643

Specificity: of the 440 clean claims, 35 were wrongly flagged, so 405/440=0.920405/440 = 0.920.

92% specificity sounds strong until you apply it at a realistic base rate. At 3% hallucination in production, a run of 10,000 claims contains 300 hallucinations and 9,700 clean claims. The detector flags 0.75×300=2250.75 \times 300 = 225 true ones and 0.08×9700=7760.08 \times 9700 = 776 false ones. Precision in production is 225/1001=22.5%225/1001 = 22.5\% — more than three quarters of your alerts are wrong, and your reviewers will stop reading them within a fortnight. A detector validated at one base rate and deployed at another will disappoint you in exactly this way.

What this means when you ship something

The deployment question is never "does this model hallucinate" — every model does. It is "what is the cost of one hallucination here, and what does the system do about it before the user sees it".

Cost of one errorDetectable by the user?Required design
Catastrophic (medical, legal, financial advice)NoDo not automate the claim. Model drafts, qualified human signs, and the signature is the product.
Severe (published content, external comms)SometimesGrounded generation with mandatory citations; every claim resolved; human review before publication
Moderate (internal search, support drafts)UsuallyRetrieval grounding, inline sources, visible confidence, easy correction path
Low (brainstorming, first drafts, code suggestions)Yes, quicklyShip it; the user's next action verifies it

Three concrete practices follow from that table, and they are worth more than any single detection technique.

Make the highest-risk token classes structurally verifiable. Numbers, dates, names and citations cause most of the damage and are the easiest to check mechanically. Require the model to emit them with a span reference into the source, then verify the span exists and contains the value. This converts an open-ended judgement problem into a string lookup.

Give the model somewhere to put uncertainty. The lawyer's model said the cases were real partly because there was no path in the conversation to any other answer. A response format with an explicit insufficient_evidence branch, and a system prompt that rewards using it, measurably reduces fabrication — and, more importantly, makes the fabrication that remains visible in your logs.

Log claim-level verdicts in production, not just in evaluation. Offline hallucination rates go stale the moment your retrieval corpus, your prompt, or the model version changes. Running the cheap checks — citation resolution, numeric span verification, NLI against the retrieved context — on live traffic gives you a rate that is current, and gives you the sample of real failures from which the next round of checks gets written.