Course Content
Python for AI and Data Science
5 sections · 13 lessons
Working with Datasets — Loading, Splitting, and Preparing Data for Scikit-learn
A team builds a model to predict which customers will cancel next month. They split the data randomly, train, and measure an AUC of 0.94. They deploy it. In production it scores 0.55 — barely better than a coin toss.
The dataset covered two years. A random split scattered those two years across both halves, so the model was trained on data from November and tested on data from the previous March. It had learned patterns from the future and been graded on the past. In production, the future is not available, and everything it relied on was gone.
The model was fine. The split was wrong. How you divide your data determines what your evaluation number means, and a number that means nothing is worse than no number at all, because you act on it.
Getting data in
| Source | Call | Good for |
|---|---|---|
| Built-in datasets | load_iris(), load_diabetes(), load_wine() | Testing code; too small and too clean for anything else |
| Downloadable | fetch_openml("credit-g"), fetch_california_housing() | Realistic size and messiness |
| Synthetic | make_classification(), make_regression(), make_blobs() | Data with known properties, to test whether your code works |
| Your own | pd.read_csv(), pd.read_parquet() | Everything real |
Synthetic generators are more useful than they look, because you control the truth. If you build a pipeline and want to know whether it can detect a signal at all, generate data where you know there are exactly 5 informative features among 20:
1from sklearn.datasets import make_classification23X, y = make_classification(4 n_samples=5000, n_features=20,5 n_informative=5, n_redundant=5, n_repeated=0,6 weights=[0.9, 0.1], # deliberately imbalanced7 flip_y=0.02, # 2% label noise, like real data8 random_state=42,9)If your feature-selection step cannot recover those five, the problem is your code, not your data.
Why split at all
A model with enough capacity can memorise its training data perfectly and learn nothing generalisable. Watch it happen:
1from sklearn.tree import DecisionTreeClassifier23deep = DecisionTreeClassifier(random_state=42) # no depth limit4deep.fit(X_train, y_train)56print(f"train {deep.score(X_train, y_train):.3f}") # 1.0007print(f"test {deep.score(X_test, y_test):.3f}") # 0.901A perfect training score is not a success, it is a symptom. The tree grew until every training row sat in its own leaf; it has stored the answers rather than learned the pattern. Only the held-out score is informative, and the gap between the two is your overfitting measurement.
Training accuracy tells you how well the model memorised. It is not a result and should never be reported as one.
1from sklearn.model_selection import train_test_split23X_train, X_test, y_train, y_test = train_test_split(4 X, y,5 test_size=0.2, # 20% held out6 random_state=42, # reproducible: same split every run7 stratify=y, # preserve class proportions8 shuffle=True,9)Why stratify matters more as data gets rarer
Suppose 3% of 1,000 rows are positive — 30 cases. A random 20% test set expects 6 of them, but the actual number follows a binomial distribution and can easily land at 2 or 11. With 2 positives in your test set, your recall can only take the values 0, 0.5 or 1, so your metric jumps around wildly for reasons that have nothing to do with the model.
1import numpy as np23X_demo = np.zeros((1000, 3)) # the features do not matter here4y_demo = np.array([1] * 30 + [0] * 970) # 3% positive56for seed in range(5):7 _, _, _, yt = train_test_split(X_demo, y_demo, test_size=0.2, random_state=seed)8 print(f"seed {seed}: {yt.mean():.1%} positive in test")9# without stratify: 2.5%, 3.0%, 2.0%, 2.5%, 1.0% <- your score depends on luck10# with stratify: 3.0%, 3.0%, 3.0%, 3.0%, 3.0%Use stratify=y on every classification split. There is no downside.
Two splits or three
| Set | Share | Used for | How often you may look |
|---|---|---|---|
| Training | 60–80% | Fitting parameters | Constantly |
| Validation | 10–20% | Choosing models and hyperparameters | Many times |
| Test | 10–20% | The final, honest estimate | Once |
1X_temp, X_test, y_temp, y_test = train_test_split(2 X, y, test_size=0.2, random_state=42, stratify=y)3X_train, X_val, y_train, y_val = train_test_split(4 X_temp, y_temp, test_size=0.25, random_state=42, stratify=y_temp)5# 0.25 of the remaining 80% = 20% of the originalThe reason for a separate validation set is subtle and important. If you try forty configurations and pick the one that scores best on the test set, that score is now optimistic — you selected for it. The test set has quietly become part of your training procedure. Keeping a set you look at exactly once, at the end, is the only way to get a number you can put in front of someone.
Splits where shuffling is the bug
| Your data | Random split gives | Use instead |
|---|---|---|
| Ordered in time | Training on the future | TimeSeriesSplit, or split at a date |
| Several rows per customer | Same customer in both halves | GroupKFold, GroupShuffleSplit |
| Near-duplicate rows | Test rows that are copies of training rows | Deduplicate before splitting |
| Nested structure — patients, sites, sessions | Leakage across the group boundary | Split by the outermost unit |
1from sklearn.model_selection import TimeSeriesSplit, GroupShuffleSplit23# Time: every fold trains on the past and tests on the future4for train_idx, test_idx in TimeSeriesSplit(n_splits=5).split(X_sorted_by_date):5 ...67# Groups: a customer_id appears in training or in test, never both8splitter = GroupShuffleSplit(test_size=0.2, n_splits=1, random_state=42)9train_idx, test_idx = next(splitter.split(X, y, groups=df.customer_id))The group case is easy to miss because the rows genuinely are different — different transactions, different visits, different sessions. But if customer 4471 has 300 transactions split across both halves, the model can identify that customer from their features and recall their behaviour. It looks like generalisation and it is recall.
The rule that governs all preprocessing
Any number learned from data — a mean, a standard deviation, a category list, a median for imputation — must be learned from the training set alone.
1from sklearn.preprocessing import StandardScaler23# WRONG: the scaler sees the test set's distribution4X_scaled = StandardScaler().fit_transform(X)5X_train, X_test = train_test_split(X_scaled, ...)67# RIGHT8X_train, X_test, y_train, y_test = train_test_split(X, y, ...)9scaler = StandardScaler().fit(X_train)10X_train_s = scaler.transform(X_train)11X_test_s = scaler.transform(X_test)The wrong version usually inflates your score by only a little — one or two points — which is exactly why it survives. It is not large enough to look suspicious, and it makes your results better, so nobody investigates. Then production performance comes in below the reported figure and nobody can explain why.
Preprocessing transformers
Scaling numeric features
| Transformer | Formula | Result range | Use when |
|---|---|---|---|
StandardScaler | (x−μ)/σ | Mean 0, sd 1 | The default; roughly symmetric data |
MinMaxScaler | (x−min)/(max−min) | 0 to 1 | Bounded inputs, images, neural nets |
RobustScaler | (x−median)/IQR | Unbounded | Outliers are present and real |
The difference matters most when there is an extreme value. Take salaries of 30k, 35k, 40k, 45k and 5,000k. MinMaxScaler maps that top value to 1.0 and squashes the other four into the range 0 to 0.003 — they become indistinguishable, and four fifths of your data has been destroyed by one row. RobustScaler uses the median and interquartile range, both of which ignore the extreme, so the four ordinary salaries stay spread out.
Which models need scaling at all:
| Needs scaling | Does not care |
|---|---|
| k-nearest neighbours, SVM, k-means — they measure distance | Decision trees, random forests, gradient boosting — they split on thresholds |
| Linear and logistic regression with regularisation — the penalty is scale-dependent | Naive Bayes |
| Neural networks — gradients behave badly otherwise |
Encoding categories
1from sklearn.preprocessing import OneHotEncoder, OrdinalEncoder23# Unordered categories: one column per level4enc = OneHotEncoder(handle_unknown="ignore", sparse_output=False, drop="first")56# Genuinely ordered categories: one column, order supplied by you7ord_enc = OrdinalEncoder(categories=[["small", "medium", "large"]])Using an ordinal encoder on unordered data is a real and common mistake. Encoding red=0, green=1, blue=2 tells a linear model that green sits halfway between red and blue and that blue is twice green, which is nonsense and will show up as strange coefficients.
handle_unknown="ignore" is not optional in production. Training data contains the categories it contains; the first live request with a new city name crashes an encoder that was not told what to do about it.
For high-cardinality columns — postcodes, product IDs, thousands of levels — one-hot encoding produces thousands of nearly-empty columns. Group rare levels into "Other", or use a target-based encoding fitted strictly inside cross-validation folds.
Imputation
1from sklearn.impute import SimpleImputer23num_imputer = SimpleImputer(strategy="median", add_indicator=True)4cat_imputer = SimpleImputer(strategy="most_frequent")add_indicator=True appends a binary column recording where the value was missing. It costs one column and preserves the signal in the missingness itself, which is often genuinely predictive — someone who declines to give their income is telling you something.
Pipelines and mixed column types
Real tables have numeric and categorical columns needing different treatment. ColumnTransformer routes each group to its own pipeline, and the whole thing behaves as one estimator:
1from sklearn.compose import ColumnTransformer2from sklearn.pipeline import Pipeline3from sklearn.preprocessing import StandardScaler, OneHotEncoder4from sklearn.impute import SimpleImputer5from sklearn.ensemble import RandomForestClassifier67numeric = ["age", "income", "tenure"]8categorical = ["region", "plan", "channel"]910numeric_pipe = Pipeline([11 ("impute", SimpleImputer(strategy="median", add_indicator=True)),12 ("scale", StandardScaler()),13])1415categorical_pipe = Pipeline([16 ("impute", SimpleImputer(strategy="constant", fill_value="Unknown")),17 ("encode", OneHotEncoder(handle_unknown="ignore", sparse_output=False)),18])1920preprocess = ColumnTransformer([21 ("num", numeric_pipe, numeric),22 ("cat", categorical_pipe, categorical),23], remainder="drop")2425model = Pipeline([26 ("prep", preprocess),27 ("clf", RandomForestClassifier(n_estimators=300, random_state=42)),28])2930model.fit(X_train, y_train)31print(model.score(X_test, y_test))Everything this buys you is structural rather than cosmetic. Leakage becomes difficult: every fit happens on training data by construction, including inside every cross-validation fold. The object serialises to a single file, so what you deploy is the exact preprocessing plus the exact model, not a model plus a page of notes about how to prepare the inputs. And when a new row arrives with a missing value and an unseen category, it is handled identically to how it was handled in training.
If your preprocessing is not in a pipeline, sooner or later training and serving will disagree about how to prepare a row, and the failure will be silent.
Imbalanced classes
With 2% positives, accuracy is worthless — always predicting the majority scores 98%. Three responses, in increasing order of intervention:
| Approach | How | Note |
|---|---|---|
| Change the metric | Precision, recall, F1, average precision | Do this first; it is free and always correct |
| Weight the classes | class_weight="balanced" | Built in, no data invented, usually enough |
| Resample | SMOTE, random undersampling | Training data only — never the test set |
RandomForestClassifier(class_weight="balanced", random_state=42)The rule about resampling is absolute. If you oversample before splitting, synthetic copies of training rows land in your test set and you are evaluating on data derived from what you trained on. Your reported recall becomes a fiction. Resampling belongs inside the training fold, which is another argument for putting it in a pipeline.
The checklist before you fit anything
| Check | Code |
|---|---|
| No missing target | y.isna().sum() == 0 |
| Split is not random when it must not be | Time-ordered? Grouped? Deduplicated? |
No identifier columns left in X | X.nunique() == len(X) finds them |
| Class balance recorded | y.value_counts(normalize=True) |
| Train and test look alike | Compare X_train.describe() with X_test.describe() |
| Nothing fitted outside a pipeline | Search your file for fit_transform on test data |
random_state set everywhere | Otherwise your results are not reproducible |
The most valuable line on that list is the last-but-one. Before you spend a week on model architecture, spend ten minutes comparing the summary statistics of your training and test sets. If a feature's mean differs sharply between them, your split was not random with respect to that feature — and whatever caused that will cause your production data to differ too.
Nearly all of the difference between a model that survives contact with reality and one that does not is decided here, before any algorithm is chosen. The split defines what your score means; the pipeline defines whether that score was honestly earned.