Course Content
Evaluating and Testing GenAI Models
4 sections · 13 lessons
Evaluating Creativity, Coherence, and Factuality
A marketing team asks a model for product descriptions. They run the outputs past two reviewers with the instruction "rate the quality from 1 to 5". Reviewer one gives an average of 4.3; reviewer two gives 2.8. The team argues for a week about which reviewer is right.
Neither is. Reviewer one was rating whether the copy was engaging. Reviewer two was rating whether the claims about the product were true. One description said the jacket was "waterproof to 20,000mm" — a specification nobody had ever given the model. It read beautifully and it was invented. Both reviewers scored it accurately on the dimension they had in their head, and the single 1-to-5 scale silently averaged two different constructs into a number that meant nothing.
That is the core problem with evaluating open-ended generation. There is no single axis. A response can be true and dull, vivid and false, elegant and self-contradicting. Any evaluation that collapses these into one score will produce disagreements that look like rater unreliability but are really construct confusion. The fix is to separate the axes, define each one operationally, and score them independently.
Three axes that are genuinely different
| Dimension | The question | Ground truth lives in | Fails as |
|---|---|---|---|
| Factuality | Are the claims true, or supported by the provided source? | The world, or the source document | Fabrication, wrong numbers, false attribution |
| Coherence | Does the text hold together as one consistent argument? | The text itself, internally | Contradiction, non-sequitur, topic drift, dangling reference |
| Creativity | Is it novel, varied, and non-obvious while still fitting the brief? | The distribution of other plausible outputs | Cliché, repetition, mode collapse, or novelty that is just incoherence |
They are not independent, and the dependencies are the interesting part. Two matter most:
- Factuality and creativity pull in opposite directions. The mechanism that lets a model produce a surprising phrase is the same mechanism that lets it produce a surprising fact. Turn the temperature up and both rise.
- Coherence is a precondition, not a peer. An incoherent text cannot be meaningfully scored for factuality, because you cannot even determine what it is claiming.
Here is that first trade-off measured on 200 product descriptions from the same model at different sampling temperatures:
| Temperature | Factuality (claim-level) | Distinct-2 (lexical diversity) | Human "engaging" mean |
|---|---|---|---|
| 0.2 | 94.1% | 0.38 | 2.9 |
| 0.7 | 89.6% | 0.61 | 3.8 |
| 1.0 | 81.3% | 0.73 | 4.0 |
| 1.3 | 68.7% | 0.79 | 3.1 |
Read the last row carefully. At temperature 1.3 diversity is still climbing, but engagement has fallen, because the text has started to wander. Diversity metrics keep rewarding what humans have already begun to reject. A single-number evaluation would have picked 1.3 or 1.0; the multi-dimensional one shows that 0.7 is the sensible operating point and that going past it buys nothing but hallucination.
Any evaluation that reports one number for open-ended generation has already made a weighting decision on your behalf, and it has not told you what weights it used.
Evaluating factuality
The first decision is which kind of factuality you mean, because the methods differ completely.
| Kind | Definition | Example task | How to check |
|---|---|---|---|
| Faithfulness (grounded) | Every claim is supported by the provided source | Summarisation, RAG answering | Entailment against the source — fully decidable |
| Factual accuracy (open-world) | Every claim is true of the world | Open question answering | Retrieval plus verification, or a knowledge base |
| Internal consistency | The text does not contradict itself | Any long-form output | Pairwise contradiction detection within the text |
Faithfulness is the tractable one and should be your default when the task provides a source. You are not asking "is this true", which is hard, but "does this follow from that", which a natural-language-inference model can decide.
Claim decomposition: the step everyone skips
Scoring a whole paragraph as "factual: yes/no" throws away almost all the information and produces terrible agreement between raters. Decompose into atomic claims first — each a single, independently checkable proposition.
Output: "The Ravenwood jacket is waterproof to 20,000mm, weighs 340g, and was designed by the team behind the Alpine Pro line."Atomic claims: C1 The Ravenwood jacket is waterproof to 20,000mm. C2 The Ravenwood jacket weighs 340g. C3 The Ravenwood jacket was designed by the Alpine Pro team.Verdicts against the product spec sheet: C1 UNSUPPORTED (spec sheet says 10,000mm) -> CONTRADICTED C2 SUPPORTED (spec sheet: 340g) C3 NOT_MENTIONED in source -> UNSUPPORTEDNow the metrics are well defined. Over 50 outputs yielding 213 claims, of which 189 are supported:
And separate out the contradictions, because they are far more damaging than mere omissions from the source. If 6 of the 24 unsupported claims actively contradict the source:
Report both. A model with 88.7% faithfulness and 0% contradictions is adding harmless colour; one with 88.7% faithfulness and 8% contradictions is actively lying about your product, and the headline number is identical.
Verification methods, and when each is right
| Method | How | Cost per claim | Good for | Fails when |
|---|---|---|---|---|
| Reference comparison | Check claim against a known-correct answer | Near zero | Closed tasks with one right answer | Multiple valid answers exist |
| NLI entailment | Run a trained entailment model on (source, claim) | ~0.0001 USD | Grounded summarisation and RAG at scale | Claims needing multi-sentence or numeric reasoning |
| Retrieval + verification | Search for evidence, then judge the claim against what you find | ~0.01 USD | Open-world facts | The retriever misses the evidence; absence is read as falsehood |
| LLM verifier with sources | Prompt a strong model with claim + evidence, force a citation | ~0.003 USD | Nuanced or compound claims | Judge shares the generator's blind spots |
| Expert human | Domain specialist checks | 1–15 USD | Medical, legal, financial; validating the above | Cannot scale; use on samples |
The dangerous failure is the third row's caveat. If your retriever fails to find evidence for a true claim and you score "no evidence" as "false", you will systematically punish correct-but-obscure statements. Use three verdicts — SUPPORTED, CONTRADICTED, INSUFFICIENT_EVIDENCE — and report the third as its own number rather than folding it into either side.
A factuality rubric people can actually apply
| Score | Anchor |
|---|---|
| 5 | Every claim verifiable and correct; uncertainty explicitly flagged where it exists |
| 4 | All substantive claims correct; a minor detail is imprecise but not misleading |
| 3 | Main claims correct; one peripheral claim unsupported or wrong |
| 2 | A central claim is wrong, or several peripheral ones are |
| 1 | Fabricated entities, invented citations, or a confidently stated falsehood on the main point |
The anchors do the work. "Rate factuality 1–5" produces kappa around 0.3; the table above routinely produces 0.7, because raters are now answering the same question.
Evaluating coherence
Coherence has three separable components, and lumping them together is why coherence ratings are so noisy.
- Local cohesion — do adjacent sentences connect? Do pronouns and definite references resolve?
- Global structure — is there an arc? Does the ending answer the beginning?
- Logical consistency — does anything contradict anything else?
Measuring local cohesion with embeddings
Embed each sentence and compute the cosine similarity of every adjacent pair. Coherent prose shows a moderate, stable value; a topic break shows a sharp dip.
1import numpy as np2from sentence_transformers import SentenceTransformer34encoder = SentenceTransformer("all-mpnet-base-v2")56def cohesion_profile(sentences):7 emb = encoder.encode(sentences, normalize_embeddings=True)8 sims = [float(np.dot(emb[i], emb[i + 1])) for i in range(len(emb) - 1)]9 return {10 "mean_adjacent": float(np.mean(sims)),11 "min_adjacent": float(np.min(sims)),12 "break_index": int(np.argmin(sims)), # where the text jumps13 "profile": sims,14 }A worked reading. Seven sentences give adjacent similarities [0.71, 0.65, 0.68, 0.18, 0.63, 0.70]. The mean is:
A mean of 0.59 looks fine. The minimum of 0.18 at position 4 is the actual finding: the text changes subject between sentence 4 and sentence 5 and never comes back. Report the minimum and the location, not just the mean — coherence failures are local events that averages hide.
Be careful in the other direction too. Very high adjacent similarity (say every pair above 0.9) usually means the model is restating the same sentence repeatedly. Cohesion has an optimum, not a maximum.
Structural checks that are worth automating
- Dangling references. Any definite noun phrase or pronoun in sentence n whose referent never appears in sentences 1…n. "The aforementioned constraint" with no prior constraint is a hard error, cheaply detected.
- Entity continuity. Track which entities appear in which sentences. Coherent text reuses a small entity set with a pattern of transitions; incoherent text introduces a new named entity in almost every sentence and never mentions it again.
- Discourse markers without antecedents. "However", "therefore", "as a result" opening a sentence that does not contrast with or follow from the previous one.
Contradiction detection
Run an NLI model over all pairs of sentences in the output and flag any pair labelled contradiction with high confidence. This is O(n2) in sentences, which is fine for a 20-sentence output (190 pairs) and needs windowing beyond a few hundred.
1from itertools import combinations2from transformers import pipeline34nli = pipeline("text-classification", model="roberta-large-mnli", top_k=None)56def contradictions(sentences, threshold=0.85):7 hits = []8 for i, j in combinations(range(len(sentences)), 2):9 scores = nli({"text": sentences[i], "text_pair": sentences[j]})10 best = max(scores, key=lambda s: s["score"])11 if best["label"] == "CONTRADICTION" and best["score"] >= threshold:12 hits.append((i, j, round(best["score"], 3)))13 return hitsThe threshold matters enormously. At 0.5 you will flag every pair of sentences that merely discuss different things; at 0.95 you will miss genuine contradictions phrased indirectly. Tune it on a labelled set of 100 outputs and report the precision you achieved at your chosen threshold, so readers know what a flag is worth.
Evaluating creativity
Creativity is the hardest of the three because it is defined relative to a distribution rather than a ground truth. Break it into three measurable things.
Diversity: is the model producing varied outputs?
Distinct-n is the ratio of unique n-grams to total n-grams across a set of generations. Across 100 product descriptions containing 4,200 tokens with 1,050 distinct unigrams and 4,100 bigrams of which 2,870 are distinct:
The trap: distinct-n falls mechanically as total length grows, because the denominator grows without bound while vocabulary saturates. Never compare distinct-n between sets of different total length. Either truncate every set to the same token count or use the expectation-adjusted variant that normalises for length.
Self-BLEU measures how much each generation resembles the others: compute BLEU for each output using the remaining outputs as references, then average. High self-BLEU means the model is producing the same text with different nouns. A model with self-BLEU 0.62 is much closer to mode collapse than one at 0.21.
Novelty: is it different from the training data and the prompt?
Two failure modes hide behind a good diversity score: memorisation (the output is copied from training data) and prompt echoing (the output is a rearrangement of the input).
Measure the longest verbatim n-gram shared with the source or with a reference corpus. If a 25-word span of a "creative" story appears verbatim in a known corpus, that is retrieval, not creation. A practical threshold: flag any shared span of 12 or more tokens for review.
Surprisal: is it non-obvious?
Score the output under a separate reference language model and look at the per-token surprisal distribution. Formulaic text sits at low surprisal throughout. Genuinely inventive text has a heavier right tail — a few high-surprisal tokens where the unexpected choice was made — while remaining coherent elsewhere. Text that is high-surprisal everywhere is not creative; it is broken.
Creativity metrics can only ever tell you an output is unusual. Whether unusual is good is a question only the coherence and factuality scores, read alongside, can answer.
A creativity rubric
| Score | Anchor |
|---|---|
| 5 | An angle a competent professional would not have thought of, executed cleanly, on-brief |
| 4 | Fresh phrasing and a non-obvious structure; recognisably above template output |
| 3 | Competent and correct but predictable; any practitioner would produce something similar |
| 2 | Formulaic; recognisable stock phrases; interchangeable with other outputs |
| 1 | Repetitive or degenerate, or "original" only by being off-brief or nonsensical |
Combining the dimensions without averaging away the truth
The instinct is to take a weighted mean. Consider two outputs on a 1–5 scale with weights 0.4 factuality, 0.3 coherence, 0.3 creativity:
| Output | Factuality | Coherence | Creativity | Weighted mean |
|---|---|---|---|---|
| A | 5 | 3 | 3 | 0.4(5)+0.3(3)+0.3(3)=3.8 |
| B | 1 | 5 | 5 | 0.4(1)+0.3(5)+0.3(5)=3.4 |
Output B scores 3.4 — only 0.4 below A, and comfortably above a "ship if above 3.0" bar. But B has a factuality score of 1, meaning it contains a confidently stated falsehood about your product. No amount of elegance compensates. The weighted mean has laundered a hard failure into a mid-range number.
The correct structure is gates then scores:
1GATES = {"factuality": 3, "safety": 5} # hard minimums23def verdict(scores: dict) -> dict:4 failed = [d for d, floor in GATES.items() if scores.get(d, 0) < floor]5 if failed:6 return {"decision": "REJECT", "reason": f"gate failed: {failed}"}7 weights = {"coherence": 0.4, "creativity": 0.35, "style_fit": 0.25}8 quality = sum(scores[d] * w for d, w in weights.items())9 return {"decision": "ACCEPT", "quality": round(quality, 2)}Factuality and safety are gates: below the floor, nothing else is considered. Coherence, creativity and style are scores: they rank the outputs that survived. This mirrors how a human editor actually works, and it stops the arithmetic from doing something no editor would.
Building a rubric that holds up
A rubric is only as good as the agreement it produces. Six steps, in order.
- Name the dimensions from observed failures, not from theory. Collect 50 bad outputs, write down what is wrong with each, and cluster. The clusters are your dimensions. Rubrics designed in a meeting always contain one dimension nobody can score.
- Write each dimension as a question with a yes/no or a countable answer wherever possible. "Does every numeric claim match the spec sheet?" beats "Is it accurate?"
- Anchor every scale point with an example drawn from your real outputs. A scale with anchors at only 1 and 5 will have all its mass at 3 and 4.
- Pilot on 30 items with two raters and compute agreement before spending the annotation budget.
- Fix the rubric where raters disagreed, not the raters. Disagreement is nearly always a rubric defect.
- Re-measure agreement, then freeze and version the rubric. A rubric edited mid-run makes the first half of your data incomparable with the second.
Worked agreement calculation on an ordinal scale
For 1–5 ratings, Cohen's kappa is wrong because it treats a 4-versus-5 disagreement as identical to a 1-versus-5. Use the intraclass correlation. Take five items rated by two raters:
| Item | Rater 1 | Rater 2 | Item mean |
|---|---|---|---|
| 1 | 4 | 5 | 4.5 |
| 2 | 2 | 2 | 2.0 |
| 3 | 5 | 4 | 4.5 |
| 4 | 3 | 3 | 3.0 |
| 5 | 1 | 2 | 1.5 |
| Rater mean | 3.0 | 3.2 | Grand mean 3.1 |
Between-items sum of squares: 2×[(1.4)2+(−1.1)2+(1.4)2+(−0.1)2+(−1.6)2]=2×7.70=15.40, so with 4 degrees of freedom MSR=3.85.
Between-raters: 5×[(−0.1)2+(0.1)2]=0.10, with 1 degree of freedom, so MSC=0.10.
Total sum of squares about the grand mean is 16.90, so the residual is 16.90−15.40−0.10=1.40 over 4 degrees of freedom, giving MSE=0.35.
An ICC of 0.85 is good agreement: the raters are tracking the same underlying quality and their small disagreements are noise around it. Below about 0.60, stop annotating and rewrite the rubric — you are buying data you cannot use.
What this means when you evaluate a real generation system
Start by writing down what a bad output looks like on your specific task, in your own words, from real examples. That list is your dimension set, and it will not match the generic factuality/coherence/creativity triple exactly — a legal drafting tool cares about citation validity and clause completeness; a children's story generator cares about age-appropriateness and narrative arc. The three axes here are a starting frame, not a taxonomy to adopt wholesale.
Then split each dimension by cost. Anything decidable by rule — numeric claims against a spec, required fields present, longest shared span with the source, dangling references — becomes an automatic check that runs on every generation. Anything requiring an entailment decision goes to an NLI model or an LLM verifier, sampled across all traffic. Only the genuinely irreducible judgements — is this engaging, is this on-brand, is this the sort of thing our best writer would have produced — consume human hours, on a sample sized so the confidence intervals are usable.
And keep the dimensions separated all the way to the report. The moment they are averaged into a single "quality score", you have recreated the argument the marketing team had: two people looking at the same number, each seeing a different construct, with no way to tell who is right.