Course Content
Evaluating and Testing GenAI Models
4 sections · 13 lessons
Common Hallucination Types (Fabricated Facts, Reasoning Errors)
In 2023 a federal court in New York sanctioned two lawyers who had filed a brief citing six judicial decisions that did not exist. The citations were not garbled or misremembered. They had case names, reporter volumes, page numbers, courts, years, and quoted passages — all in the correct format, all fabricated. When the opposing counsel could not find them, one of the lawyers asked the model whether the cases were real. It said yes.
What makes this the canonical example is not the embarrassment; it is the shape of the failure. The output was not degraded, garbled, or obviously uncertain. It was maximally fluent and maximally confident precisely where it was maximally wrong. That inversion — confidence uncorrelated with correctness — is the defining property of hallucination and the reason it cannot be caught by the intuitions that catch ordinary software bugs.
To measure hallucination you first have to be able to name what kind you are looking at, because the detection method, the metric, and the fix are different for each kind. A single "hallucination rate" that pools all of them together tells you nothing actionable.
What a hallucination is, precisely
A hallucination is content the model presents as fact that is either unsupported by the provided source or false about the world, generated with no signal distinguishing it from content that is supported or true.
Three parts of that definition earn their place:
- Presented as fact. A model writing fiction on request is not hallucinating. The failure requires an implicit assertion of truth.
- Unsupported or false. These are different failures, and conflating them is the most common measurement error in the field. More on this below.
- No distinguishing signal. This is what makes it hard. The model does not encode "I am now making this up". Token probabilities on a fabricated citation often look much like those on a real one.
Why it happens at all
Not mysticism — four concrete mechanisms:
| Mechanism | What it does | Symptom it produces |
|---|---|---|
| Next-token training objective | The model is trained to produce plausible continuations, never verified ones. Nothing in the loss distinguishes true from true-sounding. | Fluent fabrication in the exact style of correct answers |
| Parametric compression | Facts are stored as distributed weights, not as retrievable records. Rare facts are stored approximately, and approximate storage of a name or a date reconstructs as a different name or date, not as a blank. | Entity confusion; plausible wrong numbers |
| Pattern completion pressure | Having emitted "According to a study published in", the model is now under enormous distributional pressure to emit a journal name. There is no token for "actually I do not know one". | Fabricated citations, invented statistics |
| Alignment for helpfulness | Preference training rewards answers users rated highly. Users rate confident, complete answers above hedged or refusing ones. | Overconfidence; reluctance to say "I do not know" |
A model has no internal representation of the difference between remembering and inventing. Both are the same operation — sampling a plausible continuation — and only one of them happens to be right.
The taxonomy that matters: intrinsic versus extrinsic
The primary split is defined relative to a source document, if the task provides one.
| Intrinsic hallucination | Extrinsic hallucination | |
|---|---|---|
| Definition | The output contradicts the source | The output makes a claim the source does not address |
| Example | Source says revenue rose 12%; summary says revenue rose 21% | Source says revenue rose 12%; summary adds "driven by strong European demand" (not in source) |
| Detectable by | Entailment against the source alone — fully decidable | Entailment gives "neutral"; you need external verification to know if it is true |
| Always an error? | Yes, unconditionally | No — it may be true and useful, or true and out of scope, or false |
| Typical fix | Better grounding, constrained decoding, copy mechanisms | Explicit scope instructions; citation requirements |
This distinction is not academic. Intrinsic hallucinations can be detected automatically at scale with an NLI model and no internet access. Extrinsic ones cannot — they require retrieval, and retrieval failure is indistinguishable from falsehood unless you handle it deliberately. Pooling them into one rate makes the metric un-actionable, because half of it is cheap to fix and half is not.
When there is no source document — open-ended question answering — the analogous split is closed-domain (the model contradicts something in its context or instructions) versus open-domain (the model asserts something false about the world).
The failure types you will actually see
Fabricated entities and citations
The model invents a source, paper, product, API, person, or legal case with correct-looking metadata. This is the most dangerous type because the format is the credibility signal, and the format is perfect.
Prompt: "Cite a study on the effect of sleep on memory consolidation."Output: "Ratcliffe & Moreno (2019), 'Sleep-Dependent Consolidation of Declarative Memory', Journal of Cognitive Neuroscience 31(4), pp. 612-628, found a 34% improvement in recall."Reality: The journal exists. The volume/issue numbering is plausible. The authors, title, page range and finding do not exist.Detection is unusually easy for this one if you build for it: resolve every citation against a real index (Crossref, PubMed, a package registry, a case database). A citation that does not resolve is a hard failure with no judgement call required. Any system that emits references and does not do this is choosing not to catch its most catastrophic error class.
Named entity substitution
The model produces the right kind of thing with the wrong identity: the correct role attributed to a different person, a real quotation attributed to the wrong speaker, the right event in the wrong city. This comes directly from the compression mechanism — nearby entities in embedding space are interchangeable under approximate recall.
Numeric and statistical fabrication
Numbers are the highest-risk tokens in any output. They are short, they carry disproportionate meaning, they cannot be sanity-checked by reading fluency, and models produce them with total confidence. "Roughly 40% of respondents" appearing in a summary of a source that reported 27% is invisible to every fluency-based check and fatal to a decision made on it.
Reasoning errors dressed as conclusions
Distinct from factual fabrication: every premise is true, and the inference is invalid.
Premises (all true): All members of the committee are economists. Dr Okafor is an economist.Stated conclusion: Therefore Dr Okafor is on the committee.Common sub-forms: affirming the consequent (as above), unit errors in arithmetic chains, off-by-one in date arithmetic, and confusing correlation with causation while citing a correct study. Fact-checking each claim individually passes all of them, because each individual claim is true. Only checking the inference catches it — which is why claim-level verification alone gives a falsely reassuring picture on reasoning-heavy tasks.
False attribution
Real content, wrong origin: a genuine finding credited to the wrong study, a real quote to the wrong author, a real API method to the wrong library. Especially insidious in technical documentation, where pandas.DataFrame.explode_all() reads exactly like something that ought to exist.
Anachronism and temporal error
Claims that are true of one time asserted about another: the model states current officeholders, prices, versions, or population figures from its training data as though they hold now, or places an event before something it depended on. Any deployment answering time-sensitive questions needs an explicit temporal check, not just a factual one.
Source unfaithfulness in grounded tasks
Given a document and asked to summarise, the model adds information from its parametric memory. This is often true information — which is what makes it hard to argue about — but it is out of scope, unverifiable by the reader against the document, and destroys the guarantee that makes grounded generation useful in the first place.
Partial truth
The hardest type to score. The claim is 80% right with a wrong detail embedded: correct mechanism, wrong dosage; correct API, wrong argument order; correct historical event, wrong year. Binary claim-level labelling forces raters into an arbitrary call, and the resulting inter-rater agreement is poor. Handle it by decomposing further — "correct mechanism" and "correct dosage" are two claims, and scoring them separately eliminates the ambiguity.
What it costs you, by domain
| Domain | Typical hallucination | Cost of one instance | Detectability |
|---|---|---|---|
| Software documentation | Non-existent method or flag | Low — fails immediately, developer investigates | High: execute it |
| Legal | Fabricated case citation | Severe — sanctions, malpractice | High if you resolve citations; otherwise nil |
| Medical | Wrong dosage, invented contraindication | Catastrophic and irreversible | Low — reads exactly like correct guidance |
| Financial | Invented figure in a summary of a filing | Severe — decisions and disclosure risk | Medium: numbers are checkable against the source |
| Customer support | Invented policy or entitlement | Moderate — but you may be held to it | Medium: check against the policy corpus |
| Creative writing | Invented detail | None — it is the task | Not applicable |
Note the pattern in column four: the domains where hallucination is most costly are the ones where it is least detectable by a non-expert reader. That inverse relationship is the whole argument for building verification into the system rather than relying on the user to notice.
Quantifying hallucination correctly
The base metric is claim-level. Decompose each response into atomic claims, label each, and count.
Worked example: 50 responses yield 200 atomic claims. 23 are unsupported.
The clustering correction that almost everyone omits
Those 200 claims are not 200 independent observations. They come from 50 responses, about 4 claims each, and claims within a response are correlated — a response that goes off the rails tends to go off the rails several times. Treating them as independent understates your uncertainty.
The naive standard error:
The design effect for clustered sampling, with average cluster size m=4 and intra-cluster correlation ρ=0.3:
Your 95% interval on the hallucination rate is therefore 11.5%±6.1%, or roughly [5.4%, 17.6%] — not the [7.1%, 15.9%] the naive calculation suggests. A competing model at 8.9% is comfortably inside that interval and has not been shown to be better.
Claims are clustered inside responses. Ignore that and you will report differences as real that are entirely within the noise of which fifty responses you happened to sample.
The practical shortcut, if estimating ρ feels like overreach: bootstrap by response, not by claim. Resample whole responses with replacement, recompute the rate, and read the percentiles. It handles the clustering without requiring you to estimate anything.
1import numpy as np23def hallucination_rate_ci(responses, iters=2000, seed=0):4 """responses: list of lists of 0/1 labels, one list per response.5 1 = hallucinated claim. Bootstraps at the RESPONSE level."""6 rng = np.random.default_rng(seed)7 flat = [c for r in responses for c in r]8 point = sum(flat) / len(flat)9 n = len(responses)10 draws = []11 for _ in range(iters):12 idx = rng.integers(0, n, n)13 sampled = [c for i in idx for c in responses[i]]14 draws.append(sum(sampled) / len(sampled))15 lo, hi = np.percentile(draws, [2.5, 97.5])16 return round(point, 4), round(lo, 4), round(hi, 4)Report it broken down, never pooled
A single 11.5% figure hides everything you need to act on. The useful table:
| Type | Count | Rate | Severity | Fixable by |
|---|---|---|---|---|
| Intrinsic (contradicts source) | 7 | 3.5% | High | Grounding, decoding constraints |
| Extrinsic, verified false | 4 | 2.0% | High | Retrieval + refusal on low evidence |
| Extrinsic, unverifiable | 9 | 4.5% | Medium | Scope instructions, citation requirement |
| Reasoning error | 3 | 1.5% | High | Decomposition, verification of steps |
Now the engineering priority is obvious: intrinsic contradictions are the largest high-severity bucket and the cheapest to detect, so they go first. The pooled 11.5% would not have told you that.
Detecting them: red flags and real methods
Fast heuristics for a human reviewer or a cheap pre-filter:
- Any specific number, date, percentage, or measurement — the highest-risk token class
- Any proper noun that could be looked up: person, paper, case, product, API symbol
- Superlatives and absolutes: "the first", "the only", "always", "never"
- Confident answers to questions containing a false presupposition ("Why does aspirin cure tinnitus?")
- Suspiciously round statistics: "approximately 40% of users" with no source
- Answers to questions about events after the model's training cutoff
Real methods, with their trade-offs:
| Method | How it works | Catches | Cost | Limitation |
|---|---|---|---|---|
| NLI against source | Entailment model on (source, claim) | Intrinsic only | Very low | Silent on anything the source omits |
| Citation resolution | Look every reference up in a real index | Fabricated sources | Very low | Only applies to referenced claims |
| Retrieval + verification | Search, then judge the claim against retrieved evidence | Open-domain falsehood | Medium | Retrieval misses read as falsehood |
| Self-consistency sampling | Sample the same prompt k times; measure agreement across samples | Low-confidence fabrication | k× generation cost | Consistently-held false beliefs pass cleanly |
| Token-probability signals | Flag spans with low mean token probability or high entropy | Some uncertainty | Near zero | Weak correlation; needs logit access |
| Execution / tooling | Run the code, call the API, evaluate the arithmetic | Everything decidable | Low | Only for executable claims |
Self-consistency deserves a caution. It measures whether the model is stable, not whether it is right. If the model has a firmly encoded false belief — a common misconception that appeared thousands of times in training — it will produce the same wrong answer at every temperature, and self-consistency will pronounce it reliable. Stability and truth are different properties, and this method only sees the first.
Evaluating your detector, not just your model
A detector needs its own numbers. Suppose across 500 claims, 60 are genuinely hallucinated. Your detector flags 80, of which 45 are true hallucinations.
Specificity: of the 440 clean claims, 35 were wrongly flagged, so 405/440=0.920.
92% specificity sounds strong until you apply it at a realistic base rate. At 3% hallucination in production, a run of 10,000 claims contains 300 hallucinations and 9,700 clean claims. The detector flags 0.75×300=225 true ones and 0.08×9700=776 false ones. Precision in production is 225/1001=22.5% — more than three quarters of your alerts are wrong, and your reviewers will stop reading them within a fortnight. A detector validated at one base rate and deployed at another will disappoint you in exactly this way.
What this means when you ship something
The deployment question is never "does this model hallucinate" — every model does. It is "what is the cost of one hallucination here, and what does the system do about it before the user sees it".
| Cost of one error | Detectable by the user? | Required design |
|---|---|---|
| Catastrophic (medical, legal, financial advice) | No | Do not automate the claim. Model drafts, qualified human signs, and the signature is the product. |
| Severe (published content, external comms) | Sometimes | Grounded generation with mandatory citations; every claim resolved; human review before publication |
| Moderate (internal search, support drafts) | Usually | Retrieval grounding, inline sources, visible confidence, easy correction path |
| Low (brainstorming, first drafts, code suggestions) | Yes, quickly | Ship it; the user's next action verifies it |
Three concrete practices follow from that table, and they are worth more than any single detection technique.
Make the highest-risk token classes structurally verifiable. Numbers, dates, names and citations cause most of the damage and are the easiest to check mechanically. Require the model to emit them with a span reference into the source, then verify the span exists and contains the value. This converts an open-ended judgement problem into a string lookup.
Give the model somewhere to put uncertainty. The lawyer's model said the cases were real partly because there was no path in the conversation to any other answer. A response format with an explicit insufficient_evidence branch, and a system prompt that rewards using it, measurably reduces fabrication — and, more importantly, makes the fabrication that remains visible in your logs.
Log claim-level verdicts in production, not just in evaluation. Offline hallucination rates go stale the moment your retrieval corpus, your prompt, or the model version changes. Running the cheap checks — citation resolution, numeric span verification, NLI against the retrieved context — on live traffic gives you a rate that is current, and gives you the sample of real failures from which the next round of checks gets written.