Evaluating and Testing GenAI Models

Building Automated Evaluation Dashboards


At 08:40 on a Tuesday the accuracy tile on a support-summarisation dashboard went from a green 95.2% to an amber 91.8%. Six people joined a call. They rolled back the previous night's prompt change, paged the retrieval team, and spent three hours reading transcripts.

The tile was right about the arithmetic and wrong about everything else. Monday's 95.2% was computed over 4,000 scored items. Tuesday's 91.8% was computed over 180, because the sampling job that feeds the evaluator had been failing silently since midnight. Nobody had regressed anything. The dashboard had simply divided a smaller numerator by a smaller denominator and painted the result amber.

That same week the dashboard missed a real regression: summaries had started dropping customer order numbers, costing the support team about forty minutes a day. No tile moved, because no tile measured it.

Both failures come from one mistake: treating a dashboard as a display rather than as a measuring instrument. A display shows numbers. An instrument has a known precision, a known failure mode, and a documented range over which you are entitled to believe it. Everything below is about building the second thing.

The amber tile that paged six people94.194.695.995.291.89594.40123456n = 4,000amber,n = 180On 180 items the tile's interval is roughly plus or minus 4 points, wider than the drop it reported.
Every tile is an estimate, and one without its interval turns ordinary sampling noise into a three-hour incident.

The four questions a dashboard has to answer

Dashboards go wrong when they are built as a pile of charts rather than as answers to specific decisions. Four questions are worth building for, and each demands a different cadence, sample size and alerting rule.

QuestionDecision it drivesCadenceWhat it needs
Is it broken right now?Roll back, page someoneMinutesCheap signals computed on live traffic: error rate, refusal rate, latency percentiles, output-length collapse
Is it drifting?Investigate this weekDailyA fixed frozen eval set, scored identically every day, plus a smoothing rule that tolerates daily noise
Did the change help?Ship or revert a candidatePer releasePaired scoring of both versions on identical items, with a confidence interval
Where is it broken?Prioritise the fixOn demandPer-item rows you can filter, group and read — not aggregates

The Tuesday incident was a question-one alarm firing on a question-two number. Live-traffic health signals are computed on whatever traffic arrived and are allowed to be volatile; a quality metric on a frozen eval set should have a stable denominator, and a change in that denominator is itself an alertable event. Mixing the two into a single tile is the most common structural error in evaluation dashboards.

Every tile must belong to exactly one of those four questions. A tile that serves two of them will be tuned for neither and will mislead on both.

The number on the tile is an estimate, and it wobbles

Take the Tuesday numbers seriously for a moment and ask whether a drop from 95.2% to 91.8% would have been alarming even if the sampling job had been healthy.

An accuracy of 95.2% measured on 4,000 items is a sample proportion, and its standard error — the typical amount it would move if you drew a fresh sample of the same size — is:

SE=p^(1−p^)n=0.952×0.0484000=1.1424×10−5=0.0034SE = \sqrt{\frac{\hat{p}(1-\hat{p})}{n}} = \sqrt{\frac{0.952 \times 0.048}{4000}} = \sqrt{1.1424 \times 10^{-5}} = 0.0034

0.34 percentage points. That is a well-measured number. Tuesday's 91.8% on 180 items:

SE=0.918×0.082180=4.182×10−4=0.0205SE = \sqrt{\frac{0.918 \times 0.082}{180}} = \sqrt{4.182 \times 10^{-4}} = 0.0205

2.05 percentage points — six times as wobbly. The difference between the two days has standard error:

SEdiff=1.1424×10−5+4.182×10−4=4.296×10−4=0.0207SE_{\text{diff}} = \sqrt{1.1424 \times 10^{-5} + 4.182 \times 10^{-4}} = \sqrt{4.296 \times 10^{-4}} = 0.0207

The observed drop of 3.4 points is therefore z=0.034/0.0207=1.64z = 0.034 / 0.0207 = 1.64 standard errors from zero, which corresponds to p≈0.10p \approx 0.10. The 95% interval for the true change runs from 0.034−1.96(0.0207)=−0.0070.034 - 1.96(0.0207) = -0.007 to 0.034+1.96(0.0207)=+0.0750.034 + 1.96(0.0207) = +0.075: somewhere between slightly better and 7.5 points worse. Three hours of six engineers' time were spent on a number that could not distinguish those two worlds.

How big does the daily sample need to be?

Turn the question round. If you want a genuine 2-point drop from a 95% baseline to register as a clean three-standard-error event, you need SE≤2/3=0.667SE \le 2/3 = 0.667 points, so:

n=p^(1−p^)SE2=0.95×0.050.006672=0.04754.449×10−5≈1068n = \frac{\hat{p}(1-\hat{p})}{SE^2} = \frac{0.95 \times 0.05}{0.00667^2} = \frac{0.0475}{4.449 \times 10^{-5}} \approx 1068

About 1,100 scored items per day. That single number should drive the design of your scoring job. If you can only afford to score 200 items a day, then your daily tile cannot resolve a 2-point move and you must either accept a slower cadence — a rolling seven-day window gives you 1,400 items — or accept that the tile only catches catastrophes.

Put the uncertainty on the tile itself

The fix is not a footnote or a tooltip. It is the tile's primary rendering. Three rules, all cheap:

  • Show nn next to every aggregate. Had the tile read "91.8% (n = 180)" beside Monday's "95.2% (n = 4,000)", the broken sampling job would have been obvious in two seconds instead of three hours.
  • Show the interval, at the same font size as the value. "91.8% ± 4.0" tells the reader immediately how much to believe.
  • Give the tile a fourth state. Green, amber and red are not enough. Add grey: insufficient data, triggered whenever nn falls below the number your sizing calculation demands. A grey tile is an honest tile; an amber one computed on 180 items is a lie with a colour attached.

Alert arithmetic: why dashboards cry wolf

Alert fatigue is usually described as a cultural problem. It is an arithmetic problem, and it is predictable before you ship.

Suppose you monitor 12 metrics, check each one hourly against a three-standard-deviation control limit, and everything is perfectly healthy. Under a normal approximation, the probability that any single healthy check breaches a two-sided 3σ limit is 0.0027. Over a week:

12 metrics×24 hours×7 days=2016 checks12 \text{ metrics} \times 24 \text{ hours} \times 7 \text{ days} = 2016 \text{ checks}

E[false alarms]=2016×0.0027=5.4 per week\mathbb{E}[\text{false alarms}] = 2016 \times 0.0027 = 5.4 \text{ per week}

Five and a half pages a week for nothing at all. Add four model variants to the comparison view and you are checking 8,064 times a week, giving 21.8 false alarms — three a day, every day, all of them spurious. Within a fortnight the channel is muted, and the one real alarm arrives into a muted channel.

Two fixes, with their costs stated honestly:

FixArithmeticExpected false alarmsWhat it costs you
Do nothing (3σ, hourly, 12 metrics)2016×0.00272016 \times 0.00275.4 / weekCredibility
Raise the limitSet per-check α=1/2016=4.96×10−4\alpha = 1/2016 = 4.96\times10^{-4}, which is z=3.48z = 3.481.0 / weekSmall real shifts stop firing at all
Require two consecutive breaches0.00272×2016=0.01470.0027^2 \times 2016 = 0.01471 every 68 weeksOne check of delay; halves sensitivity at the margin

The consecutive-breach rule is usually the right trade, and its sensitivity cost is bounded. A shift landing exactly on the 3σ limit breaches any given check with probability 0.5, so a pair fires with probability 0.25. But a 4σ shift — an actual break — breaches each check with probability 0.84, so two consecutive breaches occur immediately with probability 0.71 and the alarm fires within four checks with probability 0.93. You give up almost nothing on real failures and remove 99.7% of the false ones.

Before you add a metric to the alerting set, multiply its check frequency by its false-positive rate. If the product is more than about one page per week, you are not adding monitoring; you are subtracting attention.

Severity should map to an action, not to a number

A severity scheme only works if each level names what happens next. Otherwise everything becomes "warning" and nothing gets done.

LevelTriggerRoutingRequired response
CriticalTwo consecutive breaches of a user-visible metric, or any breach of a safety metricPage on-callRoll back or mitigate within 30 minutes
WarningSmoothed trend crosses its limit; a slice regresses while the aggregate holdsTeam channel, working hoursTriage ticket, owner assigned same day
Data-qualityDenominator changes by more than 30%, scoring job late, judge parse-failure rate risesTeam channelFix the pipeline before believing any tile

The third row is the one nobody builds, and the one that would have saved Tuesday morning. Your evaluation pipeline is itself a production system.

Choosing what goes on the dashboard

The instinct is to display everything measurable. Resist it: every tile competes for the same fixed attention, and a screen with thirty numbers carries about as much information as one with none. A workable split is four headline metrics with alerts, plus a second tier that exists for drill-down and never pages anyone.

MetricQuestion it answersTypical targetAlert atCadence
Task accuracy / acceptabilityDid the model do the job?0.95Two consecutive days below 0.92Daily, frozen set
Hallucination rateIs it inventing things?< 0.05Smoothed rate above 0.08Daily, frozen set
Latency p95Is it usable?< 1,500 msp95 above 2,500 ms for 10 minLive, per minute
User satisfactionDo people accept the output?4.5 / 5Weekly mean below 4.0Weekly, live feedback
Second tier — visible on drill-down, never paged: coherence, relevance, output length distribution, refusal rate, cost per request, retrieval hit rate, judge swap-consistency, annotator agreement

The composite score trap

A tempting move is to collapse the second tier into one weighted "quality score" and put that on the headline tile. Watch what it does. Take weights of 0.3 for factuality, 0.2 each for coherence, language quality and relevance, and 0.1 for diversity, on a 1–5 scale.

DimensionWeightModel AModel B
Factuality0.34.53.5
Coherence0.24.55.0
Language quality0.24.55.0
Relevance0.24.55.0
Diversity0.14.55.0
Composite4.504.55

Model B's composite is higher — 0.3(3.5)+0.2(5.0)+0.2(5.0)+0.2(5.0)+0.1(5.0)=1.05+1.0+1.0+1.0+0.5=4.550.3(3.5) + 0.2(5.0) + 0.2(5.0) + 0.2(5.0) + 0.1(5.0) = 1.05 + 1.0 + 1.0 + 1.0 + 0.5 = 4.55 — while being a full point worse on the only dimension that can get you sued. Weighted composites let strength on cheap dimensions buy off weakness on expensive ones, and fluency is always the cheapest dimension to improve.

If you want one headline number, gate it instead of weighting it: report the composite only when every dimension clears its own floor, and otherwise show the failing dimension. That keeps the convenience of a single tile without letting arithmetic launder a safety failure.

Latency belongs in percentiles

Mean latency is the least useful number on the screen. A service with a 580 ms mean, a 2,400 ms p95 and a 7,100 ms p99 serving two million requests a day leaves 20,000 requests — 1% of two million — waiting more than seven seconds, and the mean never shows it.

And percentiles do not average. The mean of 24 hourly p95 values is not the daily p95, and can be wrong by a large factor in either direction. To aggregate across shards or buckets, store a histogram or t-digest per bucket, merge those, then compute the percentile from the merged sketch.

Architecture: five stages, five distinct failure modes

Text
  live traffic  ──┐                  ├──►  [1] collection ──►  [2] scoring ──►  [3] storage ──┐  frozen eval set ┘         sample,            metrics,          raw rows  │  human labels  ────────────►  tag,            judges,        + aggregates │                              buffer          heuristics                   │                                                                           ▼                        [5] alerting  ◄──────────────────────  [4] serving / API                        rules, routing,                        windowed queries,                        suppression                            n and CI attached
StageSilent failureSymptom on the dashboardGuard
CollectionSampler dies, or samples non-uniformlyTile moves for no reason; nn collapsesAlert on the denominator, not just the ratio
ScoringJudge prompt or model version changesA step change at a deploy boundaryStamp scorer version on every row; re-score a frozen reference set on upgrade
StorageOnly aggregates retainedNo drill-down possible; "why" is unanswerableKeep per-item rows with all identifiers
ServingWindow boundaries differ per chartTwo charts disagree about the same dayOne shared window definition, rendered in the header
AlertingNo suppression during known deploysStorm of alarms every releaseDeploy annotations plus a short suppression window

The storage stage deserves emphasis because it is irreversible. Aggregates cannot be un-aggregated. Store one row per scored item, and stamp it with everything that defines what the score meant:

SQL
CREATE TABLE eval_items (  id            BIGSERIAL PRIMARY KEY,  scored_at     TIMESTAMPTZ NOT NULL,  eval_set      TEXT NOT NULL,        -- 'frozen_v3' or 'live_sample'  item_id       TEXT NOT NULL,        -- stable across runs: enables pairing  slice         TEXT NOT NULL,        -- 'billing', 'technical', 'escalation'  model_name    TEXT NOT NULL,  model_version TEXT NOT NULL,        -- exact snapshot string  prompt_sha    TEXT NOT NULL,        -- hash of the system prompt  scorer        TEXT NOT NULL,        -- 'judge-v4' / 'exact-match' / 'human'  correct       BOOLEAN,  hallucinated  BOOLEAN,  latency_ms    INTEGER,  cost_micros   INTEGER,  raw_output    TEXT NOT NULL         -- you will need to read these);CREATE INDEX ON eval_items (eval_set, scored_at DESC, model_version);

With that table, the serving layer computes the value, the sample size and the interval in one query, so the API can never hand the front end a bare number:

SQL
SELECT  date_trunc('day', scored_at)            AS day,  model_version,  count(*)                                AS n,  avg(correct::int)                       AS accuracy,  1.96 * sqrt( avg(correct::int) * (1 - avg(correct::int))               / nullif(count(*), 0) )    AS ci_half_widthFROM eval_itemsWHERE eval_set = 'frozen_v3'  AND scored_at >= now() - interval '30 days'GROUP BY 1, 2ORDER BY 1 DESC;
Python
MIN_N = 1100          # from the sizing calculation abovedef tile(value, n, target, warn, crit):    """Return the payload a tile renders. Never return a bare number."""    if n < MIN_N:        return {"value": value, "n": n, "status": "grey",                "note": f"insufficient data: {n} of {MIN_N} items"}    half = 1.96 * (value * (1 - value) / n) ** 0.5    # Judge status on the interval, not the point estimate.    if value + half < crit:        status = "red"    elif value + half < warn:        status = "amber"    else:        status = "green"    return {"value": value, "n": n, "ci": half,            "status": status, "target": target}

The load-bearing line is value + half < crit. Judging status on the upper edge of the interval rather than the point estimate means a tile only turns red when the data genuinely rules out acceptable performance. That one change would have kept Tuesday green, correctly.

Charts that do not mislead

Visual design is not decoration here; a badly drawn chart produces a wrong decision as reliably as a wrong number.

SymptomCauseFix
Every wiggle looks like a crisisTruncated y-axis, e.g. 92%–96%Either include the target and the alert line in the visible range, or draw the confidence band so the wiggle sits inside it
Two charts disagree about ThursdayOne uses rolling 24 h, the other calendar days in a different timezoneOne window definition per dashboard, stated in the header
A step change nobody can explainA deploy, a prompt edit or a scorer upgradeVertical annotations for every deploy, pulled from the release log automatically
Red/green tiles unreadable for some viewersColour is the only channelPair colour with shape or text; roughly 1 in 12 men has some red–green deficiency
Chart flickers, nobody trusts itRefresh interval shorter than the aggregation windowRefresh no faster than the window; a 24 h metric refreshed every 10 seconds is theatre
Outliers smoothed into invisibilityOnly means plottedPlot the distribution or at least p50/p95; keep a "worst 20 items" panel linked to raw rows

Layer the screens. A summary screen carries the four headline tiles, an overall status, the window and the last-updated timestamp — nothing else. A trend screen shows the same metrics over time with confidence bands and deploy annotations. A comparison screen ranks candidates. A drill-down screen is a filterable table of raw rows, and it is the one that actually resolves incidents.

Drill-down, and the paradox that hides in aggregates

Here is a case that looks impossible and happens routinely. A team replaces v1 with v2, the weekly accuracy tile falls from 86% to 68%, and they roll back. The rollback was a mistake.

Slicev1 itemsv1 accuracyv2 itemsv2 accuracy
Simple lookups1,80090.0%80092.0%
Multi-step escalations20050.0%1,20052.0%
Overall2,00086.0%2,00068.0%

Check the arithmetic: v1 scored 1800(0.90)+200(0.50)=1620+100=17201800(0.90) + 200(0.50) = 1620 + 100 = 1720 correct out of 2,000, which is 86.0%. v2 scored 800(0.92)+1200(0.52)=736+624=1360800(0.92) + 1200(0.52) = 736 + 624 = 1360 out of 2,000, which is 68.0%. Yet v2 is better in every slice. The traffic mix moved: a marketing campaign pushed escalations from 10% of volume to 60%, and escalations are hard.

This is Simpson's paradox, and on a dashboard it is not a curiosity — it is a routine source of wrong rollbacks. The defence is standardisation: report accuracy reweighted to a fixed reference mix rather than to whatever arrived. Using v1's mix (90% simple, 10% escalation) as the reference:

v1=0.9(0.90)+0.1(0.50)=0.860v2=0.9(0.92)+0.1(0.52)=0.880\text{v1} = 0.9(0.90) + 0.1(0.50) = 0.860 \qquad \text{v2} = 0.9(0.92) + 0.1(0.52) = 0.880

Standardised, v2 is two points better, which is the truth. Show both numbers: the raw rate tells you what users actually experienced this week, the standardised rate tells you whether the model changed. Confusing those two is the single most expensive misreading a dashboard enables.

When an aggregate moves, always check whether the metric changed or the mix changed. They look identical on a tile and they demand opposite responses.

The comparison view is a rank-noise generator

Leaderboards invite a specific error. Five candidate models are evaluated on 300 items each and reported as 95.2, 94.1, 93.8, 92.3 and 88.9 percent. The table is sorted, the top row is declared the winner, and the ordering of the middle three is treated as meaningful.

At p^≈0.93\hat{p} \approx 0.93 and n=300n = 300, each row's standard error is:

SE=0.93×0.07300=2.17×10−4=0.0147SE = \sqrt{\frac{0.93 \times 0.07}{300}} = \sqrt{2.17 \times 10^{-4}} = 0.0147

1.47 points. The gap between two rows evaluated on independent samples has SEdiff=2×1.472=2.08SE_{\text{diff}} = \sqrt{2 \times 1.47^2} = 2.08 points, so the 95.2 versus 93.8 gap of 1.4 points is z=0.67z = 0.67 — pure noise. Even the 3-point spread from 95.2 down to 92.3 is only z=1.4z = 1.4, nowhere near conclusive. The bottom row at 88.9% is genuinely worse; everything above it is a tie that the sort order dresses up as a ranking.

Two changes make a comparison view honest. First, evaluate every model on the identical item set and analyse the paired differences — item difficulty cancels, and the required sample typically falls by a factor of three. Second, render tiers, not ranks: group models whose intervals overlap into one band labelled "indistinguishable at n = 300". A screen saying "three models are tied at the top, here are their costs and latencies" hands the decision to a criterion you can actually measure at this sample size.

What this means when you build one

Build the storage layer before a single chart. Per-item rows carrying model version, prompt hash, scorer version and slice are the asset; charts are a view over them and can be rebuilt in an afternoon. Teams that start with a charting library and keep only aggregates spend the next year unable to answer any question nobody anticipated on day one.

Size the scoring job from the arithmetic, not the budget. Work out what daily nn your smallest interesting regression needs — 1,068 items for a 2-point drop from a 95% baseline — and either fund it or widen the window until you have it. A dashboard scoring 200 items a day and alerting on daily moves is not a cheap version of a good dashboard; it is a random number generator with a colour scheme.

Validate the dashboard against a known regression before trusting it. Take a checkpoint you know is bad, or synthetically corrupt 5% of outputs, replay it, and confirm the right tile turns the right colour in the intended time. This is the only way to distinguish "no alerts because everything is healthy" from "no alerts because the scoring job has been writing nulls for a month".

Give every alert a runbook line before it may page anyone: what it means, which drill-down to open first, the two likeliest causes. An alert with no named action is a notification, and notifications get muted.

And treat the dashboard's own numbers as versioned artefacts. When you upgrade a judge, change a rubric or repair a sampler, every point before that change means something different from every point after it, so draw the boundary on the chart. The alternative is a team that spends its Tuesdays investigating its own instrumentation and its Thursdays missing the regression that actually cost the business forty minutes a day.