Python for AI and Data Science

What is Scikit-learn? The Machine Learning Toolbox


You want to find out which algorithm works best on your dataset. Reasonable plan: try five of them and compare. In a world without a shared interface, that plan looks like this:

Python
# library Amodel = KNearest(k=5)model.learn(features, labels)guesses = model.guess(new_features)# library B -- different names, and it wants the matrix transposedtree = DecisionTree()tree.build(labels, features.T)guesses = tree.classify(new_features.T)# library C -- no object at all, and it returns a tuplecoefficients, intercept = fit_linear(features, labels, regularise=0.1)guesses = features_new @ coefficients + intercept

Three algorithms, three vocabularies, three data layouts, three shapes of output. Your evaluation code has to be rewritten for each one, which means the comparison is now mostly a test of how carefully you rewrote it. Add cross-validation and hyperparameter search and the amount of glue code exceeds the amount of modelling.

Scikit-learn's central contribution is not any particular algorithm. It is the decision that every algorithm — classifier, regressor, clusterer, scaler, encoder, feature selector — exposes the same handful of methods. Once you know them, you know the whole library.

What a shared interface removesFive libraries, five APIs• A different call to train each model• A different data layout for each• Comparison code rewritten per algorithm• Swapping a model means a rewriteOne estimator contract• fit, predict,transform, score — that is it• Every estimator takesX as rows by features• One loop compares all five• Swapping a model means one line
The uniformity is the product; the algorithms are available elsewhere, the interchangeability is not.

Four methods, and that is essentially all of it

MethodMeaningFound on
fit(X, y)Learn from dataEverything (y omitted where unsupervised)
predict(X)Produce an output per rowModels
transform(X)Return a modified version of XPreprocessors
score(X, y)One-number quality measureModels

Plus two conveniences: fit_transform(X), which does both in one call, and predict_proba(X), which gives class probabilities rather than hard labels.

Python
from sklearn.datasets import load_irisfrom sklearn.model_selection import train_test_splitfrom sklearn.ensemble import RandomForestClassifierX, y = load_iris(return_X_y=True)X_train, X_test, y_train, y_test = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)model = RandomForestClassifier(n_estimators=200, random_state=42)model.fit(X_train, y_train)print(model.score(X_test, y_test))          # 0.9print(model.predict(X_test[:3]))            # [0 2 1]print(model.predict_proba(X_test[:1]))      # [[1. 0. 0.]]

Now change one word — RandomForestClassifier to LogisticRegression, or SVC, or GradientBoostingClassifier — and every other line still works. That is the payoff.

Python
from sklearn.linear_model import LogisticRegressionfrom sklearn.svm import SVCfrom sklearn.neighbors import KNeighborsClassifierfrom sklearn.ensemble import RandomForestClassifiermodels = {    "logistic":      LogisticRegression(max_iter=1000),    "svm":           SVC(),    "knn":           KNeighborsClassifier(n_neighbors=5),    "random forest": RandomForestClassifier(random_state=42),}for name, model in models.items():    model.fit(X_train, y_train)    print(f"{name:<14} {model.score(X_test, y_test):.3f}")

Eight lines to compare four fundamentally different algorithms — a nearest-neighbour lookup, a linear decision boundary, a kernel method and an ensemble of trees. Without the shared interface this is a day's work; with it, it is the thing you do before lunch on the first day.

The interface is the product. Learning one estimator teaches you every estimator, and the code you write around a model outlives your choice of model.

The data contract

Scikit-learn expects two things and is inflexible about both:

  • X — a 2D array of shape (n_samples, n_features). Rows are observations, columns are features. Always numeric.
  • y — a 1D array of length n_samples. Numbers for regression, labels for classification.
Text
        feature_0  feature_1  feature_2        yrow 0      5.1        3.5        1.4           0row 1      4.9        3.0        1.4           0row 2      6.3        3.3        6.0           2        <-------- X: (n, 3) -------->      (n,)

The single most common first error is a one-feature model:

Python
X = df["income"]              # a Series -- shape (1000,), one dimensionmodel.fit(X, y)               # ValueError: Expected a 2-dimensional container                              # but got <class 'pandas.Series'> insteadX = df[["income"]]            # double brackets -> DataFrame, shape (1000, 1)X = df["income"].values.reshape(-1, 1)    # or reshape explicitly

The error message is unusually helpful about this, but the underlying reason is worth knowing: scikit-learn cannot tell whether your 1,000 numbers are 1,000 samples of one feature or one sample of 1,000 features. The second dimension resolves the ambiguity, so it insists on having one.

Everything in X must also be numeric before it reaches an estimator. Text categories raise ValueError: could not convert string to float, and missing values raise ValueError: Input X contains NaN for most estimators. Encoding and imputation are themselves estimators, which is how they fit into the same interface.

The rule about fit that everything else depends on

fit is where learning happens, and learning must only ever happen on training data. This applies to preprocessing exactly as much as to models.

Python
from sklearn.preprocessing import StandardScalerscaler = StandardScaler()X_train_scaled = scaler.fit_transform(X_train)   # LEARN mean and std, then applyX_test_scaled = scaler.transform(X_test)         # APPLY the training values only

Calling fit_transform on the test set is the error. It looks harmless — the code runs, the shapes are right, the numbers are scaled — but it computes the test set's own mean and standard deviation, which means information about the test data has entered the process. Your reported score improves and your real-world performance does not, which is the worst possible combination because it hides itself.

fit_transform on training data. transform on everything else. If you ever type fit and a variable with test in its name on the same line, stop.

The mechanism that enforces this automatically is the Pipeline, which chains preprocessing and model into a single object with the same four methods:

Python
from sklearn.pipeline import Pipelinepipe = Pipeline([    ("scale", StandardScaler()),    ("model", LogisticRegression(max_iter=1000)),])pipe.fit(X_train, y_train)          # fits the scaler, then the modelprint(pipe.score(X_test, y_test))   # transforms with training stats, then predicts

Because the pipeline is itself an estimator, it can be cross-validated or hyperparameter-searched as one unit, and the scaler is refitted inside every fold — which is the only correct way to do it. Doing this by hand is possible and nobody does it correctly for long.

A map of the library

ModuleHoldsReach for it when
sklearn.datasetsToy and generated datasetsTesting an idea without real data
sklearn.model_selectiontrain_test_split, cross-validation, grid searchSplitting, validating, tuning
sklearn.preprocessingScalers, encoders, binningGetting features into shape
sklearn.imputeSimpleImputer, KNNImputerFilling gaps without leaking
sklearn.linear_modelLinear and logistic regression, ridge, lassoA fast, interpretable baseline
sklearn.tree, sklearn.ensembleTrees, forests, gradient boostingTabular data — usually the strongest option
sklearn.clusterK-means, DBSCAN, hierarchicalFinding groups with no labels
sklearn.decompositionPCA and friendsReducing dimensions
sklearn.metricsEvery scoring functionJudging a model properly
sklearn.pipelinePipeline, ColumnTransformerAlways, in real work

Naming conventions that let you guess

The library is consistent enough that you can often predict the name you need without looking it up.

ConventionMeansExample
Suffix ClassifierPredicts categoriesRandomForestClassifier
Suffix RegressorPredicts numbersRandomForestRegressor
Trailing underscoreAn attribute learned during fitmodel.coef_, scaler.mean_
No underscoreA setting you chosemodel.n_estimators
random_state=Makes randomness reproducibleSet it everywhere, always

The trailing underscore is a genuinely useful signal. model.coef_ does not exist until you call fit — asking for it raises AttributeError, and calling predict on an unfitted model raises NotFittedError — and its presence tells you at a glance which attributes came from the data rather than from your configuration.

Where scikit-learn is the wrong tool

Knowing the boundary saves you from fighting the library.

TaskWhy not scikit-learnUse instead
Images, audio, text generationNo deep-learning architectures, little GPU supportPyTorch, TensorFlow
Data larger than memoryEstimators expect an in-memory arrayDask, Spark, or sample it
Winning tabular competitionsHistGradientBoostingClassifier is fast and strong, but has fewer options and no GPU trainingXGBoost, LightGBM, CatBoost (same interface)
Time series forecastingNo native seasonality or trend modelsstatsmodels, Prophet, sktime
Statistical inference — p-values, confidence intervalsBuilt for prediction, not explanationstatsmodels
Causal questions — "will this intervention work?"Correlational by constructionDoWhy, EconML, an experiment

The statistical inference row surprises people. LinearRegression gives you coefficients but no standard errors and no p-values, because it was designed to predict the next value rather than to test a hypothesis about the world. If your question is "is this effect statistically significant", you want statsmodels.

Note also that XGBoost and LightGBM deliberately copy the scikit-learn interface. So does much of the wider ecosystem. Learning fit/predict/transform buys you far more than one library.

What this looks like when you actually use it

The interface turns model selection into a loop, and a loop is something you can trust and re-run:

Python
import pandas as pdfrom sklearn.pipeline import Pipelinefrom sklearn.preprocessing import StandardScalerfrom sklearn.model_selection import cross_val_scorefrom sklearn.linear_model import LogisticRegressionfrom sklearn.ensemble import RandomForestClassifier, GradientBoostingClassifierfrom sklearn.svm import SVCcandidates = {    "logistic":  LogisticRegression(max_iter=1000),    "svm":       SVC(),    "forest":    RandomForestClassifier(random_state=42),    "boosting":  GradientBoostingClassifier(random_state=42),}rows = []for name, estimator in candidates.items():    pipe = Pipeline([("scale", StandardScaler()), ("model", estimator)])    scores = cross_val_score(pipe, X_train, y_train, cv=5, scoring="f1_macro")    rows.append({"model": name, "mean": scores.mean(), "std": scores.std()})print(pd.DataFrame(rows).sort_values("mean", ascending=False).round(3))#       model   mean    std# 3  boosting  0.967  0.017# 1       svm  0.966  0.032# 0  logistic  0.958  0.027# 2    forest  0.950  0.017

Three things about that output are worth noticing. The scaler is inside the pipeline, so it is refitted within each of the five folds and never sees data it is being evaluated on. The standard deviation is reported next to the mean, because a difference of 0.001 between the top two models is far smaller than the 0.02–0.03 variation between folds — those two are tied, and picking the "winner" is picking noise. On a dataset this small, all four are within noise of each other. And the test set has not been touched at all; every number here comes from splitting the training data.

That is the practical shape of the library. You spend your effort on the parts that actually determine the outcome — the features, the split, the metric, the question — and the estimator becomes a variable you can change with one word.