Python for AI and Data Science

Working with Datasets — Loading, Splitting, and Preparing Data for Scikit-learn


A team builds a model to predict which customers will cancel next month. They split the data randomly, train, and measure an AUC of 0.94. They deploy it. In production it scores 0.55 — barely better than a coin toss.

The dataset covered two years. A random split scattered those two years across both halves, so the model was trained on data from November and tested on data from the previous March. It had learned patterns from the future and been graded on the past. In production, the future is not available, and everything it relied on was gone.

The model was fine. The split was wrong. How you divide your data determines what your evaluation number means, and a number that means nothing is worse than no number at all, because you act on it.

When shuffling is the bugJanFebMarAprMayJunJulAug01234567cut heretest =the futureA random split puts August rows beside January rows, so the model trains on months it must predict.
AUC 0.94 measured interpolation between known months; 0.55 measured the extrapolation production actually asks for.

Getting data in

SourceCallGood for
Built-in datasetsload_iris(), load_diabetes(), load_wine()Testing code; too small and too clean for anything else
Downloadablefetch_openml("credit-g"), fetch_california_housing()Realistic size and messiness
Syntheticmake_classification(), make_regression(), make_blobs()Data with known properties, to test whether your code works
Your ownpd.read_csv(), pd.read_parquet()Everything real

Synthetic generators are more useful than they look, because you control the truth. If you build a pipeline and want to know whether it can detect a signal at all, generate data where you know there are exactly 5 informative features among 20:

Python
from sklearn.datasets import make_classificationX, y = make_classification(    n_samples=5000, n_features=20,    n_informative=5, n_redundant=5, n_repeated=0,    weights=[0.9, 0.1],        # deliberately imbalanced    flip_y=0.02,               # 2% label noise, like real data    random_state=42,)

If your feature-selection step cannot recover those five, the problem is your code, not your data.

Why split at all

A model with enough capacity can memorise its training data perfectly and learn nothing generalisable. Watch it happen:

Python
from sklearn.tree import DecisionTreeClassifierdeep = DecisionTreeClassifier(random_state=42)      # no depth limitdeep.fit(X_train, y_train)print(f"train {deep.score(X_train, y_train):.3f}")   # 1.000print(f"test  {deep.score(X_test, y_test):.3f}")     # 0.901

A perfect training score is not a success, it is a symptom. The tree grew until every training row sat in its own leaf; it has stored the answers rather than learned the pattern. Only the held-out score is informative, and the gap between the two is your overfitting measurement.

Training accuracy tells you how well the model memorised. It is not a result and should never be reported as one.

Python
from sklearn.model_selection import train_test_splitX_train, X_test, y_train, y_test = train_test_split(    X, y,    test_size=0.2,          # 20% held out    random_state=42,        # reproducible: same split every run    stratify=y,             # preserve class proportions    shuffle=True,)

Why stratify matters more as data gets rarer

Suppose 3% of 1,000 rows are positive — 30 cases. A random 20% test set expects 6 of them, but the actual number follows a binomial distribution and can easily land at 2 or 11. With 2 positives in your test set, your recall can only take the values 0, 0.5 or 1, so your metric jumps around wildly for reasons that have nothing to do with the model.

Python
import numpy as npX_demo = np.zeros((1000, 3))             # the features do not matter herey_demo = np.array([1] * 30 + [0] * 970)  # 3% positivefor seed in range(5):    _, _, _, yt = train_test_split(X_demo, y_demo, test_size=0.2, random_state=seed)    print(f"seed {seed}: {yt.mean():.1%} positive in test")# without stratify: 2.5%, 3.0%, 2.0%, 2.5%, 1.0%   <- your score depends on luck# with stratify:    3.0%, 3.0%, 3.0%, 3.0%, 3.0%

Use stratify=y on every classification split. There is no downside.

Two splits or three

SetShareUsed forHow often you may look
Training60–80%Fitting parametersConstantly
Validation10–20%Choosing models and hyperparametersMany times
Test10–20%The final, honest estimateOnce
Python
X_temp, X_test, y_temp, y_test = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_val, y_train, y_val = train_test_split(    X_temp, y_temp, test_size=0.25, random_state=42, stratify=y_temp)# 0.25 of the remaining 80% = 20% of the original

The reason for a separate validation set is subtle and important. If you try forty configurations and pick the one that scores best on the test set, that score is now optimistic — you selected for it. The test set has quietly become part of your training procedure. Keeping a set you look at exactly once, at the end, is the only way to get a number you can put in front of someone.

Splits where shuffling is the bug

Your dataRandom split givesUse instead
Ordered in timeTraining on the futureTimeSeriesSplit, or split at a date
Several rows per customerSame customer in both halvesGroupKFold, GroupShuffleSplit
Near-duplicate rowsTest rows that are copies of training rowsDeduplicate before splitting
Nested structure — patients, sites, sessionsLeakage across the group boundarySplit by the outermost unit
Python
from sklearn.model_selection import TimeSeriesSplit, GroupShuffleSplit# Time: every fold trains on the past and tests on the futurefor train_idx, test_idx in TimeSeriesSplit(n_splits=5).split(X_sorted_by_date):    ...# Groups: a customer_id appears in training or in test, never bothsplitter = GroupShuffleSplit(test_size=0.2, n_splits=1, random_state=42)train_idx, test_idx = next(splitter.split(X, y, groups=df.customer_id))

The group case is easy to miss because the rows genuinely are different — different transactions, different visits, different sessions. But if customer 4471 has 300 transactions split across both halves, the model can identify that customer from their features and recall their behaviour. It looks like generalisation and it is recall.

The rule that governs all preprocessing

Any number learned from data — a mean, a standard deviation, a category list, a median for imputation — must be learned from the training set alone.

Python
from sklearn.preprocessing import StandardScaler# WRONG: the scaler sees the test set's distributionX_scaled = StandardScaler().fit_transform(X)X_train, X_test = train_test_split(X_scaled, ...)# RIGHTX_train, X_test, y_train, y_test = train_test_split(X, y, ...)scaler = StandardScaler().fit(X_train)X_train_s = scaler.transform(X_train)X_test_s = scaler.transform(X_test)

The wrong version usually inflates your score by only a little — one or two points — which is exactly why it survives. It is not large enough to look suspicious, and it makes your results better, so nobody investigates. Then production performance comes in below the reported figure and nobody can explain why.

Preprocessing transformers

Scaling numeric features

TransformerFormulaResult rangeUse when
StandardScaler(x−μ)/σ(x - \mu) / \sigmaMean 0, sd 1The default; roughly symmetric data
MinMaxScaler(x−min⁡)/(max⁡−min⁡)(x - \min)/(\max - \min)0 to 1Bounded inputs, images, neural nets
RobustScaler(x−median)/IQR(x - \text{median}) / IQRUnboundedOutliers are present and real

The difference matters most when there is an extreme value. Take salaries of 30k, 35k, 40k, 45k and 5,000k. MinMaxScaler maps that top value to 1.0 and squashes the other four into the range 0 to 0.003 — they become indistinguishable, and four fifths of your data has been destroyed by one row. RobustScaler uses the median and interquartile range, both of which ignore the extreme, so the four ordinary salaries stay spread out.

Which models need scaling at all:

Needs scalingDoes not care
k-nearest neighbours, SVM, k-means — they measure distanceDecision trees, random forests, gradient boosting — they split on thresholds
Linear and logistic regression with regularisation — the penalty is scale-dependentNaive Bayes
Neural networks — gradients behave badly otherwise

Encoding categories

Python
from sklearn.preprocessing import OneHotEncoder, OrdinalEncoder# Unordered categories: one column per levelenc = OneHotEncoder(handle_unknown="ignore", sparse_output=False, drop="first")# Genuinely ordered categories: one column, order supplied by youord_enc = OrdinalEncoder(categories=[["small", "medium", "large"]])

Using an ordinal encoder on unordered data is a real and common mistake. Encoding red=0, green=1, blue=2 tells a linear model that green sits halfway between red and blue and that blue is twice green, which is nonsense and will show up as strange coefficients.

handle_unknown="ignore" is not optional in production. Training data contains the categories it contains; the first live request with a new city name crashes an encoder that was not told what to do about it.

For high-cardinality columns — postcodes, product IDs, thousands of levels — one-hot encoding produces thousands of nearly-empty columns. Group rare levels into "Other", or use a target-based encoding fitted strictly inside cross-validation folds.

Imputation

Python
from sklearn.impute import SimpleImputernum_imputer = SimpleImputer(strategy="median", add_indicator=True)cat_imputer = SimpleImputer(strategy="most_frequent")

add_indicator=True appends a binary column recording where the value was missing. It costs one column and preserves the signal in the missingness itself, which is often genuinely predictive — someone who declines to give their income is telling you something.

Pipelines and mixed column types

Real tables have numeric and categorical columns needing different treatment. ColumnTransformer routes each group to its own pipeline, and the whole thing behaves as one estimator:

Python
from sklearn.compose import ColumnTransformerfrom sklearn.pipeline import Pipelinefrom sklearn.preprocessing import StandardScaler, OneHotEncoderfrom sklearn.impute import SimpleImputerfrom sklearn.ensemble import RandomForestClassifiernumeric = ["age", "income", "tenure"]categorical = ["region", "plan", "channel"]numeric_pipe = Pipeline([    ("impute", SimpleImputer(strategy="median", add_indicator=True)),    ("scale", StandardScaler()),])categorical_pipe = Pipeline([    ("impute", SimpleImputer(strategy="constant", fill_value="Unknown")),    ("encode", OneHotEncoder(handle_unknown="ignore", sparse_output=False)),])preprocess = ColumnTransformer([    ("num", numeric_pipe, numeric),    ("cat", categorical_pipe, categorical),], remainder="drop")model = Pipeline([    ("prep", preprocess),    ("clf", RandomForestClassifier(n_estimators=300, random_state=42)),])model.fit(X_train, y_train)print(model.score(X_test, y_test))

Everything this buys you is structural rather than cosmetic. Leakage becomes difficult: every fit happens on training data by construction, including inside every cross-validation fold. The object serialises to a single file, so what you deploy is the exact preprocessing plus the exact model, not a model plus a page of notes about how to prepare the inputs. And when a new row arrives with a missing value and an unseen category, it is handled identically to how it was handled in training.

If your preprocessing is not in a pipeline, sooner or later training and serving will disagree about how to prepare a row, and the failure will be silent.

Imbalanced classes

With 2% positives, accuracy is worthless — always predicting the majority scores 98%. Three responses, in increasing order of intervention:

ApproachHowNote
Change the metricPrecision, recall, F1, average precisionDo this first; it is free and always correct
Weight the classesclass_weight="balanced"Built in, no data invented, usually enough
ResampleSMOTE, random undersamplingTraining data only — never the test set
Python
RandomForestClassifier(class_weight="balanced", random_state=42)

The rule about resampling is absolute. If you oversample before splitting, synthetic copies of training rows land in your test set and you are evaluating on data derived from what you trained on. Your reported recall becomes a fiction. Resampling belongs inside the training fold, which is another argument for putting it in a pipeline.

The checklist before you fit anything

CheckCode
No missing targety.isna().sum() == 0
Split is not random when it must not beTime-ordered? Grouped? Deduplicated?
No identifier columns left in XX.nunique() == len(X) finds them
Class balance recordedy.value_counts(normalize=True)
Train and test look alikeCompare X_train.describe() with X_test.describe()
Nothing fitted outside a pipelineSearch your file for fit_transform on test data
random_state set everywhereOtherwise your results are not reproducible

The most valuable line on that list is the last-but-one. Before you spend a week on model architecture, spend ten minutes comparing the summary statistics of your training and test sets. If a feature's mean differs sharply between them, your split was not random with respect to that feature — and whatever caused that will cause your production data to differ too.

Nearly all of the difference between a model that survives contact with reality and one that does not is decided here, before any algorithm is chosen. The split defines what your score means; the pipeline defines whether that score was honestly earned.