Evaluating and Testing GenAI Models

Evaluating Creativity, Coherence, and Factuality


A marketing team asks a model for product descriptions. They run the outputs past two reviewers with the instruction "rate the quality from 1 to 5". Reviewer one gives an average of 4.3; reviewer two gives 2.8. The team argues for a week about which reviewer is right.

Neither is. Reviewer one was rating whether the copy was engaging. Reviewer two was rating whether the claims about the product were true. One description said the jacket was "waterproof to 20,000mm" — a specification nobody had ever given the model. It read beautifully and it was invented. Both reviewers scored it accurately on the dimension they had in their head, and the single 1-to-5 scale silently averaged two different constructs into a number that meant nothing.

That is the core problem with evaluating open-ended generation. There is no single axis. A response can be true and dull, vivid and false, elegant and self-contradicting. Any evaluation that collapses these into one score will produce disagreements that look like rater unreliability but are really construct confusion. The fix is to separate the axes, define each one operationally, and score them independently.

Three axes hiding inside one "rate it 1 to 5"Is every claim true?claim-level verificationhedge everythingDoes it hold together?cohesion,contradiction checksrepeat the topic wordIs it non-obvious?diversity, novelty, surprisalraise the temperatureThe questionMeasured byThe cheap fakeFactualityCoherenceCreativityOptimising one axis routinely costs you another.
Reviewers scoring 4.3 and 2.8 were not disagreeing about the text — they were weighting three different axes under one label.

Three axes that are genuinely different

DimensionThe questionGround truth lives inFails as
FactualityAre the claims true, or supported by the provided source?The world, or the source documentFabrication, wrong numbers, false attribution
CoherenceDoes the text hold together as one consistent argument?The text itself, internallyContradiction, non-sequitur, topic drift, dangling reference
CreativityIs it novel, varied, and non-obvious while still fitting the brief?The distribution of other plausible outputsCliché, repetition, mode collapse, or novelty that is just incoherence

They are not independent, and the dependencies are the interesting part. Two matter most:

  • Factuality and creativity pull in opposite directions. The mechanism that lets a model produce a surprising phrase is the same mechanism that lets it produce a surprising fact. Turn the temperature up and both rise.
  • Coherence is a precondition, not a peer. An incoherent text cannot be meaningfully scored for factuality, because you cannot even determine what it is claiming.

Here is that first trade-off measured on 200 product descriptions from the same model at different sampling temperatures:

TemperatureFactuality (claim-level)Distinct-2 (lexical diversity)Human "engaging" mean
0.294.1%0.382.9
0.789.6%0.613.8
1.081.3%0.734.0
1.368.7%0.793.1

Read the last row carefully. At temperature 1.3 diversity is still climbing, but engagement has fallen, because the text has started to wander. Diversity metrics keep rewarding what humans have already begun to reject. A single-number evaluation would have picked 1.3 or 1.0; the multi-dimensional one shows that 0.7 is the sensible operating point and that going past it buys nothing but hallucination.

Any evaluation that reports one number for open-ended generation has already made a weighting decision on your behalf, and it has not told you what weights it used.

Evaluating factuality

The first decision is which kind of factuality you mean, because the methods differ completely.

KindDefinitionExample taskHow to check
Faithfulness (grounded)Every claim is supported by the provided sourceSummarisation, RAG answeringEntailment against the source — fully decidable
Factual accuracy (open-world)Every claim is true of the worldOpen question answeringRetrieval plus verification, or a knowledge base
Internal consistencyThe text does not contradict itselfAny long-form outputPairwise contradiction detection within the text

Faithfulness is the tractable one and should be your default when the task provides a source. You are not asking "is this true", which is hard, but "does this follow from that", which a natural-language-inference model can decide.

Claim decomposition: the step everyone skips

Scoring a whole paragraph as "factual: yes/no" throws away almost all the information and produces terrible agreement between raters. Decompose into atomic claims first — each a single, independently checkable proposition.

Text
Output: "The Ravenwood jacket is waterproof to 20,000mm, weighs 340g,         and was designed by the team behind the Alpine Pro line."Atomic claims:  C1  The Ravenwood jacket is waterproof to 20,000mm.  C2  The Ravenwood jacket weighs 340g.  C3  The Ravenwood jacket was designed by the Alpine Pro team.Verdicts against the product spec sheet:  C1  UNSUPPORTED  (spec sheet says 10,000mm)   -> CONTRADICTED  C2  SUPPORTED    (spec sheet: 340g)  C3  NOT_MENTIONED in source                   -> UNSUPPORTED

Now the metrics are well defined. Over 50 outputs yielding 213 claims, of which 189 are supported:

Faithfulness=supported claimstotal claims=189213=0.887\text{Faithfulness} = \frac{\text{supported claims}}{\text{total claims}} = \frac{189}{213} = 0.887

And separate out the contradictions, because they are far more damaging than mere omissions from the source. If 6 of the 24 unsupported claims actively contradict the source:

Contradiction rate=6213=0.028\text{Contradiction rate} = \frac{6}{213} = 0.028

Report both. A model with 88.7% faithfulness and 0% contradictions is adding harmless colour; one with 88.7% faithfulness and 8% contradictions is actively lying about your product, and the headline number is identical.

Verification methods, and when each is right

MethodHowCost per claimGood forFails when
Reference comparisonCheck claim against a known-correct answerNear zeroClosed tasks with one right answerMultiple valid answers exist
NLI entailmentRun a trained entailment model on (source, claim)~0.0001 USDGrounded summarisation and RAG at scaleClaims needing multi-sentence or numeric reasoning
Retrieval + verificationSearch for evidence, then judge the claim against what you find~0.01 USDOpen-world factsThe retriever misses the evidence; absence is read as falsehood
LLM verifier with sourcesPrompt a strong model with claim + evidence, force a citation~0.003 USDNuanced or compound claimsJudge shares the generator's blind spots
Expert humanDomain specialist checks1–15 USDMedical, legal, financial; validating the aboveCannot scale; use on samples

The dangerous failure is the third row's caveat. If your retriever fails to find evidence for a true claim and you score "no evidence" as "false", you will systematically punish correct-but-obscure statements. Use three verdicts — SUPPORTED, CONTRADICTED, INSUFFICIENT_EVIDENCE — and report the third as its own number rather than folding it into either side.

A factuality rubric people can actually apply

ScoreAnchor
5Every claim verifiable and correct; uncertainty explicitly flagged where it exists
4All substantive claims correct; a minor detail is imprecise but not misleading
3Main claims correct; one peripheral claim unsupported or wrong
2A central claim is wrong, or several peripheral ones are
1Fabricated entities, invented citations, or a confidently stated falsehood on the main point

The anchors do the work. "Rate factuality 1–5" produces kappa around 0.3; the table above routinely produces 0.7, because raters are now answering the same question.

Evaluating coherence

Coherence has three separable components, and lumping them together is why coherence ratings are so noisy.

  • Local cohesion — do adjacent sentences connect? Do pronouns and definite references resolve?
  • Global structure — is there an arc? Does the ending answer the beginning?
  • Logical consistency — does anything contradict anything else?

Measuring local cohesion with embeddings

Embed each sentence and compute the cosine similarity of every adjacent pair. Coherent prose shows a moderate, stable value; a topic break shows a sharp dip.

Python
import numpy as npfrom sentence_transformers import SentenceTransformerencoder = SentenceTransformer("all-mpnet-base-v2")def cohesion_profile(sentences):    emb = encoder.encode(sentences, normalize_embeddings=True)    sims = [float(np.dot(emb[i], emb[i + 1])) for i in range(len(emb) - 1)]    return {        "mean_adjacent": float(np.mean(sims)),        "min_adjacent": float(np.min(sims)),        "break_index": int(np.argmin(sims)),   # where the text jumps        "profile": sims,    }

A worked reading. Seven sentences give adjacent similarities [0.71, 0.65, 0.68, 0.18, 0.63, 0.70]. The mean is:

sˉ=0.71+0.65+0.68+0.18+0.63+0.706=3.556=0.592\bar{s} = \frac{0.71+0.65+0.68+0.18+0.63+0.70}{6} = \frac{3.55}{6} = 0.592

A mean of 0.59 looks fine. The minimum of 0.18 at position 4 is the actual finding: the text changes subject between sentence 4 and sentence 5 and never comes back. Report the minimum and the location, not just the mean — coherence failures are local events that averages hide.

Be careful in the other direction too. Very high adjacent similarity (say every pair above 0.9) usually means the model is restating the same sentence repeatedly. Cohesion has an optimum, not a maximum.

Structural checks that are worth automating

  • Dangling references. Any definite noun phrase or pronoun in sentence nn whose referent never appears in sentences 1…n1 \ldots n. "The aforementioned constraint" with no prior constraint is a hard error, cheaply detected.
  • Entity continuity. Track which entities appear in which sentences. Coherent text reuses a small entity set with a pattern of transitions; incoherent text introduces a new named entity in almost every sentence and never mentions it again.
  • Discourse markers without antecedents. "However", "therefore", "as a result" opening a sentence that does not contrast with or follow from the previous one.

Contradiction detection

Run an NLI model over all pairs of sentences in the output and flag any pair labelled contradiction with high confidence. This is O(n2)O(n^2) in sentences, which is fine for a 20-sentence output (190 pairs) and needs windowing beyond a few hundred.

Python
from itertools import combinationsfrom transformers import pipelinenli = pipeline("text-classification", model="roberta-large-mnli", top_k=None)def contradictions(sentences, threshold=0.85):    hits = []    for i, j in combinations(range(len(sentences)), 2):        scores = nli({"text": sentences[i], "text_pair": sentences[j]})        best = max(scores, key=lambda s: s["score"])        if best["label"] == "CONTRADICTION" and best["score"] >= threshold:            hits.append((i, j, round(best["score"], 3)))    return hits

The threshold matters enormously. At 0.5 you will flag every pair of sentences that merely discuss different things; at 0.95 you will miss genuine contradictions phrased indirectly. Tune it on a labelled set of 100 outputs and report the precision you achieved at your chosen threshold, so readers know what a flag is worth.

Evaluating creativity

Creativity is the hardest of the three because it is defined relative to a distribution rather than a ground truth. Break it into three measurable things.

Diversity: is the model producing varied outputs?

Distinct-n is the ratio of unique n-grams to total n-grams across a set of generations. Across 100 product descriptions containing 4,200 tokens with 1,050 distinct unigrams and 4,100 bigrams of which 2,870 are distinct:

distinct-1=10504200=0.250distinct-2=28704100=0.700\text{distinct-1} = \frac{1050}{4200} = 0.250 \qquad \text{distinct-2} = \frac{2870}{4100} = 0.700

The trap: distinct-n falls mechanically as total length grows, because the denominator grows without bound while vocabulary saturates. Never compare distinct-n between sets of different total length. Either truncate every set to the same token count or use the expectation-adjusted variant that normalises for length.

Self-BLEU measures how much each generation resembles the others: compute BLEU for each output using the remaining outputs as references, then average. High self-BLEU means the model is producing the same text with different nouns. A model with self-BLEU 0.62 is much closer to mode collapse than one at 0.21.

Novelty: is it different from the training data and the prompt?

Two failure modes hide behind a good diversity score: memorisation (the output is copied from training data) and prompt echoing (the output is a rearrangement of the input).

Measure the longest verbatim n-gram shared with the source or with a reference corpus. If a 25-word span of a "creative" story appears verbatim in a known corpus, that is retrieval, not creation. A practical threshold: flag any shared span of 12 or more tokens for review.

Surprisal: is it non-obvious?

Score the output under a separate reference language model and look at the per-token surprisal distribution. Formulaic text sits at low surprisal throughout. Genuinely inventive text has a heavier right tail — a few high-surprisal tokens where the unexpected choice was made — while remaining coherent elsewhere. Text that is high-surprisal everywhere is not creative; it is broken.

Creativity metrics can only ever tell you an output is unusual. Whether unusual is good is a question only the coherence and factuality scores, read alongside, can answer.

A creativity rubric

ScoreAnchor
5An angle a competent professional would not have thought of, executed cleanly, on-brief
4Fresh phrasing and a non-obvious structure; recognisably above template output
3Competent and correct but predictable; any practitioner would produce something similar
2Formulaic; recognisable stock phrases; interchangeable with other outputs
1Repetitive or degenerate, or "original" only by being off-brief or nonsensical

Combining the dimensions without averaging away the truth

The instinct is to take a weighted mean. Consider two outputs on a 1–5 scale with weights 0.4 factuality, 0.3 coherence, 0.3 creativity:

OutputFactualityCoherenceCreativityWeighted mean
A5330.4(5)+0.3(3)+0.3(3)=3.80.4(5)+0.3(3)+0.3(3) = 3.8
B1550.4(1)+0.3(5)+0.3(5)=3.40.4(1)+0.3(5)+0.3(5) = 3.4

Output B scores 3.4 — only 0.4 below A, and comfortably above a "ship if above 3.0" bar. But B has a factuality score of 1, meaning it contains a confidently stated falsehood about your product. No amount of elegance compensates. The weighted mean has laundered a hard failure into a mid-range number.

The correct structure is gates then scores:

Python
GATES = {"factuality": 3, "safety": 5}       # hard minimumsdef verdict(scores: dict) -> dict:    failed = [d for d, floor in GATES.items() if scores.get(d, 0) < floor]    if failed:        return {"decision": "REJECT", "reason": f"gate failed: {failed}"}    weights = {"coherence": 0.4, "creativity": 0.35, "style_fit": 0.25}    quality = sum(scores[d] * w for d, w in weights.items())    return {"decision": "ACCEPT", "quality": round(quality, 2)}

Factuality and safety are gates: below the floor, nothing else is considered. Coherence, creativity and style are scores: they rank the outputs that survived. This mirrors how a human editor actually works, and it stops the arithmetic from doing something no editor would.

Building a rubric that holds up

A rubric is only as good as the agreement it produces. Six steps, in order.

  1. Name the dimensions from observed failures, not from theory. Collect 50 bad outputs, write down what is wrong with each, and cluster. The clusters are your dimensions. Rubrics designed in a meeting always contain one dimension nobody can score.
  2. Write each dimension as a question with a yes/no or a countable answer wherever possible. "Does every numeric claim match the spec sheet?" beats "Is it accurate?"
  3. Anchor every scale point with an example drawn from your real outputs. A scale with anchors at only 1 and 5 will have all its mass at 3 and 4.
  4. Pilot on 30 items with two raters and compute agreement before spending the annotation budget.
  5. Fix the rubric where raters disagreed, not the raters. Disagreement is nearly always a rubric defect.
  6. Re-measure agreement, then freeze and version the rubric. A rubric edited mid-run makes the first half of your data incomparable with the second.

Worked agreement calculation on an ordinal scale

For 1–5 ratings, Cohen's kappa is wrong because it treats a 4-versus-5 disagreement as identical to a 1-versus-5. Use the intraclass correlation. Take five items rated by two raters:

ItemRater 1Rater 2Item mean
1454.5
2222.0
3544.5
4333.0
5121.5
Rater mean3.03.2Grand mean 3.1

Between-items sum of squares: 2×[(1.4)2+(−1.1)2+(1.4)2+(−0.1)2+(−1.6)2]=2×7.70=15.402 \times [(1.4)^2 + (-1.1)^2 + (1.4)^2 + (-0.1)^2 + (-1.6)^2] = 2 \times 7.70 = 15.40, so with 4 degrees of freedom MSR=3.85MS_R = 3.85.

Between-raters: 5×[(−0.1)2+(0.1)2]=0.105 \times [(-0.1)^2 + (0.1)^2] = 0.10, with 1 degree of freedom, so MSC=0.10MS_C = 0.10.

Total sum of squares about the grand mean is 16.90, so the residual is 16.90−15.40−0.10=1.4016.90 - 15.40 - 0.10 = 1.40 over 4 degrees of freedom, giving MSE=0.35MS_E = 0.35.

ICC(2,1)=MSR−MSEMSR+(k−1)MSE+k(MSC−MSE)n=3.85−0.353.85+0.35+2(0.10−0.35)5ICC(2,1) = \frac{MS_R - MS_E}{MS_R + (k-1)MS_E + \frac{k(MS_C - MS_E)}{n}} = \frac{3.85 - 0.35}{3.85 + 0.35 + \frac{2(0.10 - 0.35)}{5}}

=3.504.20−0.10=3.504.10=0.854= \frac{3.50}{4.20 - 0.10} = \frac{3.50}{4.10} = 0.854

An ICC of 0.85 is good agreement: the raters are tracking the same underlying quality and their small disagreements are noise around it. Below about 0.60, stop annotating and rewrite the rubric — you are buying data you cannot use.

What this means when you evaluate a real generation system

Start by writing down what a bad output looks like on your specific task, in your own words, from real examples. That list is your dimension set, and it will not match the generic factuality/coherence/creativity triple exactly — a legal drafting tool cares about citation validity and clause completeness; a children's story generator cares about age-appropriateness and narrative arc. The three axes here are a starting frame, not a taxonomy to adopt wholesale.

Then split each dimension by cost. Anything decidable by rule — numeric claims against a spec, required fields present, longest shared span with the source, dangling references — becomes an automatic check that runs on every generation. Anything requiring an entailment decision goes to an NLI model or an LLM verifier, sampled across all traffic. Only the genuinely irreducible judgements — is this engaging, is this on-brand, is this the sort of thing our best writer would have produced — consume human hours, on a sample sized so the confidence intervals are usable.

And keep the dimensions separated all the way to the report. The moment they are averaged into a single "quality score", you have recreated the argument the marketing team had: two people looking at the same number, each seeing a different construct, with no way to tell who is right.