- MantraMindAI
- Blog
- AI Engineering & MLOps
When model evaluation lies: splits, leakage, and slices
Jai Rao
August 22, 202623 min read
Why a single accuracy number tells you almost nothing, how leakage and reused test sets inflate scores, and how to evaluate generative output, slices, and drift honestly.
A model gets a number on a slide — "94% accurate" — and the room nods. The questions nobody asks are the only ones that decide anything: 94% on which rows, how many times did you look at those rows before you got 94%, what does a model that always guesses the majority class score on the same rows, and how does it do on the eight percent of traffic carrying most of the business risk? Change the answer to any of those and the number moves further than the gap between the model you shipped and the model you rejected.
Evaluation is the part of machine learning where the job is to catch yourself cheating. It is adversarial, and the adversary is you: nearly every convenient shortcut in how a test set is built, a metric chosen, or a result read pushes the score in the same direction, which is up. Nobody sets out to inflate a number. An inflated number is simply what you get by not thinking carefully, which is why evaluation has to be a discipline with rules rather than a step tacked onto the end of training.
One number, and everything it refuses to tell you
Accuracy is an average, and an average summarises a distribution you chose. Your test set has some mix of classes, some mix of easy and hard cases, some mix of languages and customers and input lengths. Reweight that mix and the accuracy moves while the model's weights sit untouched. So "94% accurate" is not a property of the model. It is a property of the model paired with one particular sample of the world, and if that sample was assembled from whatever data was easy to collect, the number describes your collection process at least as much as your model.
Two habits fix most of the damage and both are cheap. First, print the trivial baseline beside every score. If the test set is 90% negatives, a function that returns False scores 90%, and a 94% model has contributed four points, not ninety-four. If a regressor gets MAE 3.1 and repeating last week's value gets 3.3, you built a pipeline to buy 0.2. Second, print the sample size and an interval. For a proportion the standard error is roughly sqrt(p * (1 - p) / n); at p = 0.94 with n = 500 that is about 0.011, so a 95% interval spans some four and a half points. Two models 1.5 points apart on 500 examples are, on your evidence, the same model. If you would rather not do algebra, bootstrap: resample the test rows with replacement a few thousand times, rescore, and read off the 2.5th and 97.5th percentiles.
A single number also cannot say what kind of wrong the model is. Two classifiers at identical accuracy can fail in completely different shapes — one spreading errors thinly over everything, the other flawless on nine tenths of the input and hopeless on the rest. Those are different products, and you only tell them apart by splitting the number, never by averaging harder.
Three splits, three different jobs
The three-way split is usually taught as a ritual and then performed as one, and the expensive failures nearly all come from mixing up the jobs.
The training set fits parameters. The validation set is where you make decisions: learning rate, depth, which of forty feature ideas survive, where the threshold goes, when to stop. It exists to be consumed. Every comparison you run on it transfers a little of its noise into your model, so after a hundred comparisons the validation score is optimistic — you selected for that specific sample. That is fine; that is the set's purpose. The test set answers exactly one question: how does this do on data that influenced none of my choices? It can only answer while it remains untouched.
What burns a test set is hard work, not cheating
You rarely burn a test set by breaking a rule. You burn it by being conscientious. Train, check the test score, notice it is low, add a feature, retrain, check again, move the threshold, check again. Forty rounds of that is hand-rolled gradient descent on the test set with your own judgement as the optimiser, and the final number is not an estimate of generalisation — it is the best of forty draws, biased upward by roughly the spread between draws.
Rules that hold up in practice: the test set opens for a release candidate, not for an experiment. Keep a plain text log of every opening, who did it and what the number was; the log's real job is to make the count visible, and a count of thirty-one is its own conclusion. If the work genuinely needs frequent honest checks, seal two sets and rotate, and when the second disagrees with the first, treat the first as spent. On small data, use k-fold cross-validation for decisions and keep one outer holdout for the final claim — and if you tune inside cross-validation, the tuning needs its own inner loop or the CV score inherits the same optimism.
Whatever structure the data has, the split must respect it. Grouped rows split by group. Ordered rows split by time. Cross-validation grants no exemption: a random 5-fold on a time series trains on the future in four folds out of five.
Leakage, the wound the field keeps giving itself
Leakage is any path by which information that will not exist at prediction time reaches the model during training. It is the most common serious bug in applied machine learning, and its signature is a result that makes you happy. That is precisely what makes it durable: the feedback loop rewards it. A bug that produces a bad number gets fixed on Tuesday; a bug that produces a great number gets put in a deck.
- Duplicated rows across splits. Scraped corpora, retried writes, augmented copies of one image, the same ticket filed twice. A random split lands a row in train and its twin in test, and the model is graded on material it memorised. Near-duplicates are the worse half, because they survive a naive de-duplication on exact equality.
- A feature computed with future information. The obvious version is a column derived from the label. The subtle version is an aggregate computed over the whole table — an average order value that includes the order you are predicting, or "days since last login" computed as of export time rather than as of the event.
- Preprocessing fit before splitting. Scalers, imputers, target encoders, vocabularies, dimensionality reduction, feature selection by correlation with the label, and oversampling. Each learns from the rows it sees, so fitting on the full table smuggles test statistics into training.
- Grouped data split at the wrong granularity. Several scans per patient, sessions per user, passages per document, frames per video. Split rows and the model learns to recognise the entity, which buys nothing in production because there the entity is new.
- A time series split at random. Training on Thursday to predict Tuesday. Anything with a trend, a season, or slow-moving user state will report a score you cannot reproduce going forward.
The grouped case is worth watching happen. This builds a dataset where each patient has a private measurement signature — a scanner quirk, a body habitus, anything stable — and a label that is a coin flip unrelated to anything learnable. No honest model can beat chance here. Standard library only, so it pastes and runs anywhere.
import randomrandom.seed(0)N_PATIENTS, SCANS = 300, 6# Each patient has a private measurement signature the scanner imprints on# every one of their scans. The label is a coin flip per PATIENT, and nothing# in the signature predicts it. Ceiling for any honest model: chance.sig = {p: [random.gauss(0, 1) for _ in range(8)] for p in range(N_PATIENTS)}label = {p: random.randint(0, 1) for p in range(N_PATIENTS)}rows = [] # (features, label, patient_id)for p in range(N_PATIENTS): for _ in range(SCANS): rows.append(([v + random.gauss(0, 0.1) for v in sig[p]], label[p], p))print(len(rows), "rows,", N_PATIENTS, "patients")Each patient contributes six scans that sit close together in feature space. Now score a nearest-neighbour classifier twice: once on a random split of rows, once with whole patients held out.
def nn_accuracy(train, test): sq = lambda a, b: sum((u - v) ** 2 for u, v in zip(a, b)) hits = sum(min(train, key=lambda r: sq(r[0], x))[1] == y for x, y, _ in test) return round(hits / len(test), 3)random.shuffle(rows) # naive: shuffle and slicecut = int(0.7 * len(rows))print("random row split :", nn_accuracy(rows[:cut], rows[cut:]))held = set(range(0, N_PATIENTS, 3)) # correct: hold out PATIENTStrain = [r for r in rows if r[2] not in held]test = [r for r in rows if r[2] in held]print("by-patient split :", nn_accuracy(train, test), "on", len(test), "scans")The output is random row split : 1.0 and by-patient split : 0.567. The first is perfect on a label that is literally a coin flip, because for every test scan there is another scan of the same patient sitting in training a hair away in feature space. The second is chance, which happens to be the truth. Nothing about the model changed between those two lines — only the question the split was asking.
The preprocessing case has a one-line fix that people skip because the wrong version is shorter to write. Here X and y are your feature matrix and labels.
from sklearn.pipeline import make_pipelinefrom sklearn.preprocessing import StandardScalerfrom sklearn.linear_model import LogisticRegressionfrom sklearn.model_selection import cross_val_score# Wrong: the scaler already saw the mean and variance of every held-out fold.X_scaled = StandardScaler().fit_transform(X)print(cross_val_score(LogisticRegression(), X_scaled, y, cv=5).mean())# Right: each fold refits the scaler on that fold's training rows only.pipe = make_pipeline(StandardScaler(), LogisticRegression())print(cross_val_score(pipe, X, y, cv=5).mean())The first version leaks the mean and variance of every held-out fold into the transform. The resulting inflation is usually small, which is exactly why it survives review: small enough to shrug at, large enough to decide which model you ship.
Catching it before it catches you
Leakage announces itself if you know the tells. A score that jumps when you add one feature and collapses when you remove it is a lead — go look at how that feature is computed. Importance dominated by something ID-shaped is a lead. A suspiciously narrow train/test gap is a lead, because real models overfit a little and one that scores identically on both often has the answer somewhere in its input. And when a score exceeds what a domain expert believes is achievable, believe the expert.
Better than tells are assertions, three lines each, living in the same job that builds the splits: no hash of a training row appears in test, the intersection of group identifiers across splits is empty, and the maximum timestamp in train is strictly less than the minimum in test. Written down that way, every leakage class above turns into a failing build instead of a discovery six months later.
Let the decision pick the metric
Before choosing a metric, answer a plain question: what does somebody do differently because of this output, and what does each kind of mistake cost them? The metric is downstream of that answer. Picking one because it is the default argument in a library call is how teams end up optimising something no user cares about.
Classification. Accuracy is honest only when classes are near-balanced and the two error types cost about the same. Otherwise the errors have to be reported separately. ROC-AUC scores how well the model ranks positives above negatives across all thresholds, which makes it useful for comparing models and useless for predicting production behaviour, because production runs exactly one threshold — so also report the numbers at the threshold you will deploy. And if the score feeds arithmetic, calibration matters more than ranking: among cases scored 0.7, roughly 70% should turn out positive. A reliability curve and a Brier score measure that. AUC is completely blind to it, because AUC only sees order, and a model whose scores are all squashed between 0.4 and 0.6 can have a beautiful AUC and be unusable in an expected-value calculation.
Regression. MAE is the typical error, in the units your users think in. RMSE is dominated by the worst few predictions, which makes it right when one large miss hurts more than several small ones and wrong when your data has outliers you have consciously decided not to care about. Report both: the ratio between them tells you how heavy the tail is. Avoid MAPE anywhere the target approaches zero, where it detonates, and remember it penalises over- and under-prediction unequally. R-squared is relative to the variance of your particular test sample, so comparing it across datasets is meaningless. Where the cost of error is genuinely asymmetric, train and evaluate with quantile loss and predict a quantile rather than a mean. And always print the naive forecast — last value, or last season's value. A great many production regressors quietly lose to it.
Ranking. Accuracy has no meaning here and position has all of it, because people do not scroll. Recall@k answers "is the right item anywhere in the candidate set I hand downstream", which is the question for a first-stage retriever. nDCG and precision@k answer "is the top of this list good", which is what the user actually experiences. MRR answers "how quickly does someone reach the one thing they came for". Choose by which of those three sentences describes your product, and report per-query spread as well as the mean, because a mean over queries hides that your head queries are excellent and your tail is noise.
| What the output decides | Metric that fits | What it still hides |
|---|---|---|
| A label, costs roughly symmetric | Accuracy against the majority-class baseline | Any subgroup where it collapses |
| A score multiplied by money | Calibration curve, Brier score | A calibrated model can still rank poorly |
| A number a person reads | MAE plus the 90th-percentile error | The handful of enormous misses |
| A number feeding a control loop | RMSE plus max error | Gets dragged by outliers you may not care about |
| Candidates for a second-stage model | Recall@k | The ordering inside k |
| A ranked list a user reads | nDCG@10, MRR, per-query spread | The gap between head and tail queries |
| Free text | Per-property checks plus a validated judge | Anything your rubric forgot to ask about |
Scoring output that has no single right answer
Generative output is where evaluation stops being a solved engineering problem. There is no key. For a summary, a rewrite, an explanation, or an answer drawn from a document, many outputs are correct, more are acceptable, and the boundary between acceptable and not is a judgement call that has to be made explicit before it can be measured.
Exact match still works where the task has a canonical form: an extracted date, an SKU, a JSON field, a choice from a fixed label set. Normalise, compare, be strict. The moment the output is a sentence, exact match reports failure for correct answers and teaches you nothing.
The usual next stop is n-gram overlap — BLEU, ROUGE and relatives — which counts shared word sequences against one or a few reference texts. These are weak for a structural reason rather than a tuning reason: a correct paraphrase sharing few words scores low, and a fluent answer with one wrong number in it scores high, because overlap cannot tell which words carry the meaning. They were designed for translation, where references are plentiful and adequacy really does track word overlap. Used honestly they have one job left, as a tripwire in a regression suite where a large drop means something changed and deserves a look. They are not a quality measure, and a move from 0.31 to 0.33 means nothing at all.
What works better is refusing to score "quality" and decomposing it into properties you can check. A summarisation task is not one judgement, it is several: does the output contain the three facts the ticket requires, does it stay under eighty words, does every named entity in it appear in the source, does it parse against the schema, does it decline when the source lacks the answer, does it avoid the four phrases legal asked about. Most of those are deterministic code. The ones that are not are still binary, which makes them cheap to label and easy to trend. A panel of eight binary rates tells you far more than one score out of ten, and when it moves you already know which property broke.
For the properties that genuinely require reading — is this explanation faithful to the source, is this tone right, is this answer even responsive — a language model can grade at a cost and speed no human panel matches. It works. It also fails in specific, repeatable ways that you have to design around rather than hope about.
- Position bias. Shown two candidates, judges systematically favour one slot over the other. The effect is comfortably large enough to flip verdicts on close pairs, which is where you most need the answer.
- Self-preference. A judge tends to score text resembling its own output more highly. Grading a model with a judge from the same family puts a thumb on the scale in the one direction you least want.
- Verbosity and format bias. Longer answers, headings and bullets beat shorter answers with the same content. If a prompt change made outputs forty percent longer, expect the judge to report an improvement your users will not feel.
- Agreeing with confident nonsense. Asked "is this correct?" with no reference material to hand, a judge grades fluency and assurance, because that is all it has. A confidently wrong answer with clean structure sails through. This is the worst failure mode, because it certifies exactly the errors you built the evaluation to catch.
- Score compression. On a one-to-ten scale nearly everything lands between six and eight, leaving a metric with about two bits of resolution and plenty of noise.
The mitigations are mechanical. Run every pairwise comparison in both orders and keep only verdicts that survive the swap; a position-dependent verdict is not evidence of anything.
FLIP = {"A": "B", "B": "A", "tie": "tie"}def compare(judge, question, a, b): """judge(question, first, second) -> "A", "B" or "tie" (which slot won). Ask twice with the slots swapped; keep only verdicts that survive it.""" first = judge(question, a, b) second = judge(question, b, a) # a is now in slot B if first == FLIP[second]: return first # stable under position swap return "tie" # position decided it -> no signalTreating inconsistent pairs as ties costs some resolution and buys a metric that does not move when you shuffle the input. Beyond that: never let the judge be the knowledge base — hand it the source and ask for a claim-by-claim verdict with a supporting quote for each, so "correct" becomes "supported by this span" instead of "sounds right". Ask for binary or pairwise judgements rather than Likert scores. Length-match the candidates, or make the rubric explicitly penalise unsupported additions. Judge with a different model family than the one under test.
Then comes the step that separates a judge from a random number generator with good manners: measure it. Hand-label 150 to 300 examples spanning easy and contested cases, and compute the judge's agreement with those labels. Now you hold an instrument with a known error rate, and can say something defensible: the judge agrees with our reviewers 88% of the time, and where it disagrees it is usually too generous about faithfulness. Re-measure whenever the judge model or its rubric changes, because that is a new instrument.
An unvalidated judge does not measure quality. It measures how much your output resembles what that judge happens to like, which is a real quantity, but not the one you promised anybody.
Human review never fully goes away, and pretending otherwise is how evaluation programmes rot. Humans define what good means in the first place, since the rubric is a human artefact and no automation invents the standard it enforces. Humans settle contested cases, own anything with legal or safety consequence, and recalibrate the judge. Budget review as a recurring cost rather than a launch task, sample deliberately instead of randomly — oversample low-confidence outputs and the slices you already distrust — and measure agreement between two independent reviewers. If reviewers disagree a third of the time, you have not found a model problem. You have found that your task definition is ambiguous, and no metric will be stable until the definition is fixed.
Slice before you celebrate
An aggregate score is a weighted average, and weighted averages are excellent at hiding things. A slice that is four percent of traffic can be at forty percent accuracy while the headline sits at ninety-one. Nobody in the room is lying. The model is simply useless for a group of people the average declined to mention.
Define the slices before you look at any results, drawn from the dimensions where failure would be a genuine problem: input language and locale, length buckets, channel or device, tenant or major account, recency, source system, the rare classes, and the hard cases you already know about. Then score each and hold it to a floor, rather than ranking them and discussing the worst.
from collections import defaultdictdef slice_report(records, key, floor=0.85, min_n=40): buckets = defaultdict(list) for r in records: buckets[r[key]].append(r["correct"]) overall = sum(r["correct"] for r in records) / len(records) lines = [] for name, hits in buckets.items(): n, acc = len(hits), sum(hits) / len(hits) flag = "THIN" if n < min_n else ("FAIL" if acc < floor else "ok ") lines.append((acc, f"{flag} {name:<7} n={n:<5} acc={acc:.3f}")) body = "\n".join(line for _, line in sorted(lines)) return f"overall n={len(records)} acc={overall:.3f} (floor={floor})\n{body}"Run over a couple of thousand graded records keyed by locale, it prints something like this.
overall n=2042 acc=0.901 (floor=0.85)FAIL hi-IN n=180 acc=0.622FAIL ja-JP n=90 acc=0.722THIN ar-EG n=22 acc=0.727ok de-DE n=250 acc=0.908ok en-US n=1200 acc=0.946ok en-GB n=300 acc=0.950Overall accuracy is 0.901 and would pass most review meetings unchallenged. Two slices sit far below the floor, and one is too thin to judge at all. Notice what the report does with the thin slice: it flags it for labelling, it does not wave it through. "We don't have enough Arabic data to measure it" is a statement about your evaluation, not evidence that the model works there.
Two disciplines make per-slice evaluation trustworthy. A floor per slice, checked in the release gate, so a regression affecting four percent of traffic can block a ship even when the average improved. And restraint about multiple comparisons: with forty slices one will look terrible by chance, so require a gap to survive a second run or a bootstrap before treating it as real. When a slice genuinely fails and you cannot fix it, narrowing the supported scope is a legitimate answer. Shipping a feature that quietly does not work for a group is not.
Evaluation as infrastructure, not an event
An evaluation that runs when somebody remembers is not a measurement system. The version that survives contact with a roadmap looks like a test suite: a fixed set of inputs, a set of expected properties, a runner, a threshold, and a build that goes red.
The suite earns its value by accretion. Every failure that reaches production becomes a permanent case in it, with the violated property written down beside it. After a year you own a file of the specific, weird, real ways your system breaks, and no change can silently reintroduce any of them. That file is worth more than any public benchmark score and it cannot be downloaded.
Pin everything that can move, because six things move and only one of them is the model: model version, prompt or feature code, preprocessing code, the eval dataset itself, the judge model and its rubric, and the sampling temperature and seed. A score that changed with no identified cause is not a result, it is an unexplained variable. Where output is non-deterministic, run each case several times, report a mean with its spread, and set the gate threshold wider than the run-to-run noise — otherwise CI flaps and people learn to ignore it, which is worse than having no gate.
Then there is the gap between the world you trained on and the world you are serving. Training data is a photograph; production is a river. The inputs can be watched immediately, before any label exists: feature histograms against the training distribution, the share of categories the model never saw, the input length distribution, distance from an embedding centroid for text, and the mix of predicted classes with the shape of the score histogram. A prediction distribution that shifts while the input distribution holds steady is usually a bug in your own pipeline. Both moving together is usually the world changing, which needs a different response.
Labels are the hard part, because in most real systems they arrive late or never. Fraud resolves in sixty days. A recommendation nobody clicked is never labelled at all. What you do have is proxies, and they are worth instrumenting deliberately: how often users edit or overwrite the output, retry and regenerate rate, escalation to a human, abandonment before completion, downstream acceptance of the suggestion, and agreement with a slower and more expensive model on a small sample. Each of these is biased and you should say so out loud: edit rate measures effort rather than correctness, and thumbs-down is dominated by the small subset of users who click things. Their value is as a trigger — a proxy that moves ten percent in a week is a reason to pull a hundred examples and read them, which is where diagnosis actually happens. And when delayed labels finally land, backfill and re-score the cohorts they belong to, never comparing a mature cohort against one still filling in: the young cohort looks better purely because its bad outcomes have not happened yet, an artefact that has fooled a lot of dashboards.
Signs your numbers are honest
An evaluation programme can be audited quickly by asking what it produces rather than what it believes. These are the artefacts worth having, and their absence is usually the whole story.
- Every reported score carries its n and an interval, and sits next to a trivial baseline.
- There is a log of how many times the test set has been opened, and somebody is uncomfortable about the count.
- The split job asserts no duplicate rows across splits, no shared group identifiers, and train timestamps strictly before test timestamps — and it fails the build, not a report.
- A per-slice table exists with a floor per slice, and thin slices are named as unmeasured instead of reported as fine.
- Generative quality is a panel of binary properties with individual rates, not one score out of ten.
- The judge has a measured agreement number, a date, and a note on which direction it errs.
- The regression suite holds a case for every production failure of the past year, and somebody can point at a build it blocked.
- Production monitoring reports input drift and at least one behavioural proxy, with a named owner who reads them.
None of this exists to produce a higher number. It exists to produce a number you would still believe after someone hostile went looking for the reason it was wrong. The most valuable instinct to build is suspicion of your own good results: when a score improves and you cannot name the change that caused it, treat it as a bug report rather than a win, and go find out what you accidentally told the model.