Evaluating and Testing GenAI Models

Mini Project: LLM Evaluation Pipeline


A team ran almost exactly this project to decide whether they could replace an expensive model with one costing a twelfth as much. They built 25 prompts, generated from both models, and got a clean-looking result: BLEU 0.31 against 0.28, and an LLM judge preferring the cheap model on 15 of the 23 items where it expressed a preference. Two numbers pointing the same way. They switched.

Six weeks later they switched back, after a support lead noticed that the cheap model was inventing plausible refund policies about once in every forty answers — a failure rate that produced roughly zero hits in a 25-item sample and several a day in production.

Nothing in their pipeline was wrong. The BLEU comparison was computed correctly, the judge prompt was reasonable, the code ran. What was missing was the arithmetic: a 15–8 preference split on 23 comparisons is 1.46 standard errors from a coin flip, and a 3-point BLEU gap on 25 paired items sits comfortably inside its own confidence interval. Their experiment did not say the cheap model was better. It said nothing at all, in a confident tone.

That is what this project is for. You will build the whole evaluation stack — automatic metrics, hallucination detection, rubric scoring, pairwise comparison, an LLM judge validated against humans — and then do the thing almost nobody does: work out what it is entitled to conclude.

The deliverable is not a winner. It is a recommendation you can defend under questioning, and "these two models are indistinguishable on our data, so choose on price" is a first-class result.

Five phases, and what each one is allowed to concludeDesign the dataset as the experimentGenerate under fixed decoding settingsAutomatic metrics, then a paired testHallucination rate with an intervalBlind human scoring and a judge
BLEU 0.31 against 0.28 on 25 prompts and 15 judge wins out of 23 are both well inside noise — the paired test is what turns a run into a decision.

What you are building

Text
  dataset.json ──► generate ──► generations.json ──┬──► automatic metrics   20-30 items    both models   paired + metadata   │    BLEU, ROUGE, BERTScore   stratified     same settings                     ├──► hallucination checks                                                    │    claims, NLI verification                                                    └──► human + hybrid                                                         rubric, pairwise, judge                     all three ──► scorecard + significance ──► recommendation
PhaseTimeArtefact you produce
1. Task, models, dataset30 mindataset.json
2. Paired generation40 mingenerations.json
3. Automatic metrics55 minlexical_scores.csv, perplexity.csv
4. Hallucination and factuality55 minhallucination_scores.csv + annotated examples
5. Human and hybrid evaluation70 minhuman_review_sheet.csv, llm_judge_results.csv
6. Scorecard and report30 minevaluation_scorecard.png + written recommendation

Install everything up front:

Bash
pip install openai anthropic bert-score sacrebleu rouge-score nltk transformers \            torch pandas matplotlib scipy scikit-learn

Phase 1 — Choose a comparison that can teach you something

Pick two models whose difference you cannot predict. A frontier model against a two-year-old 7B model tells you nothing you did not already know.

PairingExampleThe question it answers
Two vendors, similar tierA small hosted model from each of two providersWhich vendor suits this task? Exposes vendor style biases
Hosted against open-weightsA hosted small model against an 8B instruct model you run yourselfIs self-hosting viable? Quality against cost and control
Same family, two tiersThe flagship against its mini variantThe most useful in practice: is the cheap tier good enough?

The dataset is the experiment

Twenty to thirty prompts is the right size to learn the pipeline, and the wrong size to decide anything — the report must say so. Design the set to reveal as much as it can at that size:

  • Stratify deliberately. Three or four strata of roughly equal size — easy factual, multi-step reasoning, under-specified, adversarial — recorded on each item. Aggregates hide the interesting result, which is usually that the models tie everywhere except one stratum.
  • Include items with no correct answer. Four or five prompts with a false premise ("Which year did the WHO ban aspirin?") or an unknowable answer. These are the highest-yield items in the set: they separate models that say "I don't know" from models that confabulate.
  • Write references only where one is meaningful. Summarisation and factual QA need gold answers; open-ended generation has none, and inventing a reference so BLEU has something to chew on just measures similarity to your writing style.
  • Freeze it. Hash the file and do not edit it after generation begins. Quietly adding three easy items after seeing early results invalidates everything downstream.
Python
import json, hashlibdataset = [    {"id": 1, "stratum": "factual_easy",     "prompt": "In what year was the World Health Organization founded?",     "reference": "1948"},    {"id": 2, "stratum": "unanswerable",     "prompt": "Which year did the WHO ban aspirin?",     "reference": "UNANSWERABLE: the WHO has never banned aspirin."},    # ... 18-28 more, balanced across strata]blob = json.dumps(dataset, indent=2)open("dataset.json", "w").write(blob)print(len(dataset), "items; sha256:", hashlib.sha256(blob.encode()).hexdigest()[:12])

Phase 2 — Generate under controlled conditions

Every difference between the two runs that is not the model is a confound. Fix all of them: identical prompt text, temperature (use 0 or 0.3 and say which; some current reasoning models accept no temperature at all, so record that instead), maximum output length, system prompt or none at all, and the same day. Then record what you cannot control.

Python
import json, timedef run_model(call, dataset, name, version):    out = []    for item in dataset:        t0 = time.time()        text, usage = call(item["prompt"])          # your provider wrapper        out.append({**item, "model": name, "model_version": version,                    "output": text,                    "latency_ms": round((time.time() - t0) * 1000),                    "output_tokens": usage["output_tokens"],                    "cost_micros": usage["cost_micros"]})        time.sleep(0.5)                             # stay under rate limits    return outdataset = json.load(open("dataset.json"))rows = run_model(call_a, dataset, "model_a", "a-2025-06-01") + \       run_model(call_b, dataset, "model_b", "b-2025-06-01")json.dump(rows, open("generations.json", "w"), indent=2)

Store the exact version string, not the alias: aliases get repointed without notice, and the snapshot identifier is the only thing that lets you reproduce the result later. Record latency and cost per item — the recommendation depends on them and they cannot be reconstructed afterwards.

If a model refuses or returns an empty string, keep the row with an empty output. Dropping failures inflates the failing model's average, which is exactly backwards.

Phase 3 — Automatic metrics, and what they cannot settle

Compute BLEU and ROUGE against your references, then BERTScore, which compares contextual embeddings rather than exact n-grams and so credits paraphrase.

Python
import jsonimport pandas as pdfrom rouge_score import rouge_scorerfrom nltk.translate.bleu_score import sentence_bleu, SmoothingFunctionfrom bert_score import score as bert_scorerouge = rouge_scorer.RougeScorer(["rouge1", "rouge2", "rougeL"], use_stemmer=True)smooth = SmoothingFunction().method1def lexical(ref, hyp):    r = rouge.score(ref, hyp)    return {"bleu": sentence_bleu([ref.split()], hyp.split(),                                  smoothing_function=smooth),            "rouge1": r["rouge1"].fmeasure, "rougeL": r["rougeL"].fmeasure}rows = [r for r in json.load(open("generations.json")) if r.get("reference")]df = pd.DataFrame([{**r, **lexical(r["reference"], r["output"])} for r in rows])for model, g in df.groupby("model"):                      # BERTScore per model    _, _, F1 = bert_score(list(g["output"]), list(g["reference"]), lang="en")    df.loc[g.index, "bertscore_f1"] = F1.numpy()df.to_csv("lexical_scores.csv", index=False)

The paired test that decides whether the gap is real

Suppose model B averages BLEU 0.311 and model A 0.280 — a gap of 0.031. Because both models saw the same items, work with the per-item differences di=BLEUB(i)−BLEUA(i)d_i = \text{BLEU}_B(i) - \text{BLEU}_A(i), not the two means. Item difficulty then cancels, which is worth roughly a threefold reduction in required sample size and costs nothing.

Say those 25 differences have mean dˉ=0.031\bar{d} = 0.031 and standard deviation sd=0.11s_d = 0.11. The standard error of the mean difference is:

SE=sdn=0.1125=0.115=0.022SE = \frac{s_d}{\sqrt{n}} = \frac{0.11}{\sqrt{25}} = \frac{0.11}{5} = 0.022

t=dˉSE=0.0310.022=1.41(df=24),p≈0.17t = \frac{\bar{d}}{SE} = \frac{0.031}{0.022} = 1.41 \quad (df = 24), \qquad p \approx 0.17

With t0.975,24=2.064t_{0.975, 24} = 2.064, the 95% confidence interval is:

0.031±2.064×0.022=0.031±0.045=[−0.014,  +0.076]0.031 \pm 2.064 \times 0.022 = 0.031 \pm 0.045 = [-0.014,\; +0.076]

The interval contains zero. Your data is consistent with B being slightly worse and with B being substantially better. Report it as "+0.031 BLEU, 95% CI [−0.014, +0.076], not significant at n = 25" — that sentence is worth more than any number in the scorecard, because it is the only one that tells a reader how much to believe.

Python
from scipy import statspivot = df.pivot(index="id", columns="model", values="bleu").dropna()d = pivot["model_b"] - pivot["model_a"]t, p = stats.ttest_rel(pivot["model_b"], pivot["model_a"])half = stats.t.ppf(0.975, len(d) - 1) * d.std(ddof=1) / len(d) ** 0.5print(f"mean diff {d.mean():.3f}  t={t:.2f}  p={p:.3f}  "      f"95% CI [{d.mean()-half:.3f}, {d.mean()+half:.3f}]")

Perplexity: log it, never rank on it

Perplexity is the exponential of the average negative log-likelihood a model assigns to a text — roughly, how surprised it is by it. It is a fine diagnostic and a terrible ranking metric here, for two reasons. You cannot compare the two candidates' perplexities directly, because different tokenisers mean the per-token average is taken over different numbers of tokens for the same string. And if you score both models' outputs with a common third model, you measure how predictable each output is to that third model, which rewards bland, high-frequency text: "Yes, that is correct" scores beautifully and answers nothing. Log it as a fluency diagnostic and state in one sentence that it played no part in the recommendation.

Phase 4 — Hallucination and factuality

Split each output into atomic claims, verify each against the reference or a retrieved source, and report a rate.

The naive splitter — splitting on full stops — breaks on decimals ("3.5 million"), abbreviations ("Dr. Chen") and on sentences carrying two assertions. "Founded in 1948, the WHO now has 194 member states" is two claims; if one is right and one wrong, a sentence-level verdict throws away half your signal. Use an abbreviation-aware sentence splitter, then ask a model to decompose each sentence into atomic claims.

Python
DECOMPOSE = """Split the passage into atomic factual claims, one per line.Each claim must be independently checkable and must not use pronouns.Passage: {text}"""VERIFY = """Context: {context}Claim: {claim}Is the claim supported by the context? Reply with exactly one word:SUPPORTED, CONTRADICTED, or NOT_IN_CONTEXT."""def claims_and_verdicts(text, context, ask):    claims = [c.strip("- ").strip()              for c in ask(DECOMPOSE.format(text=text)).splitlines() if c.strip()]    return [(c, ask(VERIFY.format(context=context, claim=c)).strip().upper())            for c in claims]

The three-way verdict matters. Folding NOT_IN_CONTEXT into "hallucinated" punishes a model for adding true background; folding it into "fine" lets confident fabrication through. Report both rates. Cross-check a sample with a natural-language-inference model such as roberta-large-mnli, which labels a claim entailment, contradiction or neutral — an independent signal that does not share the judge's blind spots.

Is the difference in hallucination rate real?

Model A produces 6 unsupported claims out of 120, model B 11 out of 130: 5.00% against 8.46%, so B's rate looks nearly 70% higher. The standard error of the difference between two proportions is:

SE=0.050×0.950120+0.0846×0.9154130=3.958×10−4+5.957×10−4=0.0315SE = \sqrt{\frac{0.050 \times 0.950}{120} + \frac{0.0846 \times 0.9154}{130}} = \sqrt{3.958\times10^{-4} + 5.957\times10^{-4}} = 0.0315

z=0.0846−0.05000.0315=1.10,p≈0.27z = \frac{0.0846 - 0.0500}{0.0315} = 1.10, \qquad p \approx 0.27

Not close to significant — and the true uncertainty is worse, because claims cluster inside responses: one confabulated answer contributes five wrong claims at once. Treating clustered observations as independent understates the standard error, often by half again. The fix at this scale is to make the response the unit of analysis — "proportion of responses with at least one unsupported claim" — giving n=25n = 25 per model and a defensible standard error.

Whenever your nn is larger than the number of prompts you wrote, ask what is really independent. Claims inside one response are not separate experiments.

Then read the failures. Pull three or more flagged examples per model, quote them beside the source, and classify each: fabricated fact (invented entity, date or citation), entity confusion (right fact, wrong subject), reasoning error (sound premises, invalid inference), unsupported extrapolation (plausible, absent from the source). A rate tells you how often; only examples tell you what to fix.

Phase 5 — Human and hybrid evaluation

Rubric anchors, not adjectives

Rate at least three dimensions on a 1–5 scale and define each point in observable terms. A rubric saying "5 = excellent, 3 = average" measures nothing but the rater's mood.

Dimension531
FactualityEvery checkable claim is supported by the sourceOne unsupported claim, peripheral to the answerA central claim is fabricated or contradicted
CoherenceNo contradictions; each sentence follows from the lastOne confusing jump a reader can repairSelf-contradictory or unreadable
HelpfulnessFully answers the question asked, nothing extraneousAnswers partly, or buries the answer in paddingDoes not address the question

Blind scoring

Score a stratified sample of about 10 items, ideally with a second rater. Randomise which output appears first and strip all model identifiers; without that, knowing which model you expect to win produces exactly that result.

Python
import random, pandas as pddef blind_sheet(items, n=10, seed=7):    rng = random.Random(seed)    sheet, key = [], []    for it in rng.sample(items, n):        pair = [("model_a", it["a_output"]), ("model_b", it["b_output"])]        rng.shuffle(pair)        sheet.append({"id": it["id"], "prompt": it["prompt"],                      "response_1": pair[0][1], "response_2": pair[1][1]})        key.append({"id": it["id"], "response_1_is": pair[0][0]})    pd.DataFrame(sheet).to_csv("human_review_sheet.csv", index=False)    pd.DataFrame(key).to_csv("blinding_key.csv", index=False)   # open later

Agreement: why 75% agreement can mean almost nothing

Compute Cohen's kappa, not raw agreement. Suppose two raters each label 20 responses acceptable or not, agreeing "acceptable" on 11 and "not acceptable" on 4, and disagreeing on 5.

Observed agreement is po=15/20=0.75p_o = 15/20 = 0.75. But rater one called 14 of 20 acceptable and rater two 13 of 20, so agreement expected by chance alone is:

pe=1420⋅1320+620⋅720=0.455+0.105=0.560p_e = \frac{14}{20}\cdot\frac{13}{20} + \frac{6}{20}\cdot\frac{7}{20} = 0.455 + 0.105 = 0.560

κ=po−pe1−pe=0.750−0.5601−0.560=0.1900.440=0.43\kappa = \frac{p_o - p_e}{1 - p_e} = \frac{0.750 - 0.560}{1 - 0.560} = \frac{0.190}{0.440} = 0.43

Seventy-five per cent agreement sounds strong; kappa 0.43 is "moderate", and says more than half of that agreement was two people guessing the same way. Below about 0.4, rewrite the rubric anchors rather than averaging the raters together. Working alone, rate a sample twice a day apart, report self-agreement, and state the single-rater limitation.

Pairwise comparison and the LLM judge

Run the judge over the full dataset, judging every pair twice with the order swapped. Position bias is large and systematic: many judges prefer whichever response they see first, and a single-order run bakes that in silently.

Python
def judge_pair(ask, prompt, out_a, out_b):    def one(first, second):        return ask(JUDGE_PROMPT.format(prompt=prompt,                                       response_1=first, response_2=second))    fwd, rev = one(out_a, out_b), one(out_b, out_a)   # "1", "2" or "tie"    a = (fwd == "1") + (rev == "2")    b = (fwd == "2") + (rev == "1")    return {"winner": "tie" if a == b else ("A" if a > b else "B"),            "consistent": a == 2 or b == 2}

Report the swap-consistency rate: the fraction of pairs where the judge gave the same verdict in both orders. Below roughly 80% it is adding more noise than signal, and its win rate is not evidence.

Now the arithmetic the opening team skipped. Of 25 items the judge picks A on 15, B on 8, and ties on 2: a 65.2% win rate for A over 23 decisive comparisons. Under a null of no preference, A-wins is binomial with p=0.5p = 0.5, so:

SE=0.5×0.523=0.54.796=0.104,z=0.652−0.5000.104=1.46SE = \sqrt{\frac{0.5 \times 0.5}{23}} = \frac{0.5}{4.796} = 0.104, \qquad z = \frac{0.652 - 0.500}{0.104} = 1.46

That is p≈0.14p \approx 0.14. A 15–8 split is what a fair coin produces about one time in seven. How many comparisons would settle a genuine 65/35 preference at the conventional 5% level with 80% power?

n=[zα/20.25+zβp(1−p)]2(p−0.5)2=[1.96(0.5)+0.84(0.477)]20.152=1.9080.0225≈85n = \frac{\left[z_{\alpha/2}\sqrt{0.25} + z_{\beta}\sqrt{p(1-p)}\right]^2}{(p - 0.5)^2} = \frac{\left[1.96(0.5) + 0.84(0.477)\right]^2}{0.15^2} = \frac{1.908}{0.0225} \approx 85

Eighty-five decisive comparisons — about 93 items at an 8% tie rate. Your 25-item set is a quarter of what the question needs. Say so, and say what it can detect: 23 decisive comparisons give 80% power against a preference near 80/20, so it catches a rout and cannot adjudicate a close race.

Validating the judge against your humans

The judge is an instrument and needs calibrating against what it claims to approximate. On your 10 blind-scored items, compare the judge's winner with the human winner. Suppose they match on 8 of 10:

SE=0.8×0.210=0.126,95% CI=0.80±1.96(0.126)=[0.55, 1.00]SE = \sqrt{\frac{0.8 \times 0.2}{10}} = 0.126, \qquad 95\%\ \text{CI} = 0.80 \pm 1.96(0.126) = [0.55,\ 1.00]

An 80% figure from 10 items is consistent with a judge barely better than a coin flip. For a half-width of 8 points you would need n=0.16/(0.08/1.96)2≈96n = 0.16 / (0.08/1.96)^2 \approx 96 labelled items. Always quote the interval with the estimate. And check whether the disagreements are one-sided: if the judge differs from humans only where it favours the longer response, you have found length bias, which is a more useful finding than the agreement rate itself.

Phase 6 — Scorecard and the report

Build one table with every signal, its sample size and its interval. A bar chart without error bars is the visual equivalent of a mean quoted without one.

SignalModel AModel BDifference (95% CI)Verdict
BLEU (paired, n = 25)0.2800.311+0.031 [−0.014, +0.076]Inconclusive
Responses with an unsupported claim (n = 25)16%28%+12 pts, wide intervalInconclusive, worth more data
Judge win rate (23 decisive)65.2%34.8%z = 1.46, p = 0.14Inconclusive
Human factuality (n = 10)4.13.6Directionally favours AUnderpowered
Median latency1,240 ms410 msB is 3x fasterMeasured reliably
Cost per 1,000 requestsUSD 2.40USD 0.20B is 12x cheaperMeasured exactly

Notice the shape of that table, because it is the shape of most evaluations at this scale: the quality signals are inconclusive, the cost and latency signals exact. That is a finding, not a failure, and it implies a decision rule. When quality differences cannot be established and cost differs by an order of magnitude, pilot the cheap model behind a guardrail while collecting the hundred items that would settle the question.

The report must contain, in order: the setup (exact model versions, dataset composition by stratum, decoding settings); the scorecard; the classified hallucination examples; the judge validation with its interval; at least one documented disagreement between signals and your analysis of why; the power limitations, stated plainly; and the recommendation with its guardrails.

The disagreement is the most interesting page

You will find items where BLEU prefers one model and every human prefers the other. Each usual cause has a different fix:

DisagreementLikely causeWhat it tells you
High BLEU, low human factualityOutput copies the source's wording but gets a fact wrongLexical overlap is not accuracy; never gate a release on BLEU
Judge prefers B, humans prefer AB's answers are longer and better formattedVerbosity bias; test by truncating and re-judging
Low BERTScore, high human ratingA correct answer phrased unlike the referenceYour reference is one of many valid answers
Zero hallucinations, low helpfulnessThe model hedges and refusesNever report a hallucination rate without an answer rate

How the work is judged

AreaWeightWhat full marks look like
Task and data setup10%Two models justified; 20+ stratified prompts including unanswerable items; dataset frozen and hashed
Automatic metrics20%BLEU/ROUGE and BERTScore correct; perplexity logged and excluded from the decision; paired analysis, not two independent means
Hallucination and factuality20%Atomic claim decomposition; three-way verdicts; a second verifier; three or more classified examples per model
Human and hybrid evaluation30%Anchored rubric; blind randomised scoring; kappa or a documented single-rater limitation; order-swapped judging with swap-consistency; judge validated against humans, with an interval
Statistical honesty10%Every quoted difference carries an interval; required sample size computed and compared with what you had
Reporting and recommendation10%Scorecard, a documented disagreement analysed, and a recommendation whose confidence matches the evidence

Extensions, in order of value

  1. Get to 100 items. The highest-value extension by far: it moves the project from illustrative to decisive. Everything else refines an underpowered study.
  2. Cost-adjusted quality. Compute quality per unit cost and state the price at which your recommendation flips.
  3. Judge ensemble. Two or three judges voting; does ensemble-human agreement beat the best single judge, and at what cost multiple?
  4. Three-way comparison with Elo. Derive ratings from pairwise results — Elo needs far more comparisons than people expect, which is the lesson.
  5. Domain stress test. Re-run on a high-risk domain with a stricter hallucination threshold; does the ranking survive?
  6. Calibrate the judge. Add human-rated examples to the judge prompt and measure the change in agreement on held-out items.

What separates a passing project from a useful one

A passing project runs every method and produces every artefact. A useful one produces a recommendation whose confidence matches its evidence, and the difference lies almost entirely in the last hour of work.

Write the limitations before the recommendation. Sample size, single rater, references reflecting one person's writing style, a judge from the same family as one candidate, everything in one language and domain — listing these first stops the recommendation overclaiming, because you can see what it stands on.

State a decision rule, not a preference. "Pilot B on the simple-lookup stratum behind an unsupported-claim detector, keep A on escalations, revisit at 100 more labelled items" is actionable. "Model B is better" is not, and on this evidence it is not even true.

Keep the pipeline, not just the results. The dataset, generation script, scorers and analysis notebook are a reusable instrument: the next model release becomes a re-run, and every failure a human finds becomes a permanent item in the frozen set. A suite built that way gets more informative every time you use it. A one-off comparison expires the day either provider ships an update.