Course Content
Python for AI and Data Science
5 sections · 13 lessons
What is Scikit-learn? The Machine Learning Toolbox
You want to find out which algorithm works best on your dataset. Reasonable plan: try five of them and compare. In a world without a shared interface, that plan looks like this:
1# library A2model = KNearest(k=5)3model.learn(features, labels)4guesses = model.guess(new_features)56# library B -- different names, and it wants the matrix transposed7tree = DecisionTree()8tree.build(labels, features.T)9guesses = tree.classify(new_features.T)1011# library C -- no object at all, and it returns a tuple12coefficients, intercept = fit_linear(features, labels, regularise=0.1)13guesses = features_new @ coefficients + interceptThree algorithms, three vocabularies, three data layouts, three shapes of output. Your evaluation code has to be rewritten for each one, which means the comparison is now mostly a test of how carefully you rewrote it. Add cross-validation and hyperparameter search and the amount of glue code exceeds the amount of modelling.
Scikit-learn's central contribution is not any particular algorithm. It is the decision that every algorithm — classifier, regressor, clusterer, scaler, encoder, feature selector — exposes the same handful of methods. Once you know them, you know the whole library.
Four methods, and that is essentially all of it
| Method | Meaning | Found on |
|---|---|---|
fit(X, y) | Learn from data | Everything (y omitted where unsupervised) |
predict(X) | Produce an output per row | Models |
transform(X) | Return a modified version of X | Preprocessors |
score(X, y) | One-number quality measure | Models |
Plus two conveniences: fit_transform(X), which does both in one call, and predict_proba(X), which gives class probabilities rather than hard labels.
1from sklearn.datasets import load_iris2from sklearn.model_selection import train_test_split3from sklearn.ensemble import RandomForestClassifier45X, y = load_iris(return_X_y=True)6X_train, X_test, y_train, y_test = train_test_split(7 X, y, test_size=0.2, random_state=42, stratify=y8)910model = RandomForestClassifier(n_estimators=200, random_state=42)11model.fit(X_train, y_train)1213print(model.score(X_test, y_test)) # 0.914print(model.predict(X_test[:3])) # [0 2 1]15print(model.predict_proba(X_test[:1])) # [[1. 0. 0.]]Now change one word — RandomForestClassifier to LogisticRegression, or SVC, or GradientBoostingClassifier — and every other line still works. That is the payoff.
1from sklearn.linear_model import LogisticRegression2from sklearn.svm import SVC3from sklearn.neighbors import KNeighborsClassifier4from sklearn.ensemble import RandomForestClassifier56models = {7 "logistic": LogisticRegression(max_iter=1000),8 "svm": SVC(),9 "knn": KNeighborsClassifier(n_neighbors=5),10 "random forest": RandomForestClassifier(random_state=42),11}1213for name, model in models.items():14 model.fit(X_train, y_train)15 print(f"{name:<14} {model.score(X_test, y_test):.3f}")Eight lines to compare four fundamentally different algorithms — a nearest-neighbour lookup, a linear decision boundary, a kernel method and an ensemble of trees. Without the shared interface this is a day's work; with it, it is the thing you do before lunch on the first day.
The interface is the product. Learning one estimator teaches you every estimator, and the code you write around a model outlives your choice of model.
The data contract
Scikit-learn expects two things and is inflexible about both:
X— a 2D array of shape(n_samples, n_features). Rows are observations, columns are features. Always numeric.y— a 1D array of lengthn_samples. Numbers for regression, labels for classification.
feature_0 feature_1 feature_2 yrow 0 5.1 3.5 1.4 0row 1 4.9 3.0 1.4 0row 2 6.3 3.3 6.0 2 <-------- X: (n, 3) --------> (n,)The single most common first error is a one-feature model:
1X = df["income"] # a Series -- shape (1000,), one dimension2model.fit(X, y) # ValueError: Expected a 2-dimensional container3 # but got <class 'pandas.Series'> instead45X = df[["income"]] # double brackets -> DataFrame, shape (1000, 1)6X = df["income"].values.reshape(-1, 1) # or reshape explicitlyThe error message is unusually helpful about this, but the underlying reason is worth knowing: scikit-learn cannot tell whether your 1,000 numbers are 1,000 samples of one feature or one sample of 1,000 features. The second dimension resolves the ambiguity, so it insists on having one.
Everything in X must also be numeric before it reaches an estimator. Text categories raise ValueError: could not convert string to float, and missing values raise ValueError: Input X contains NaN for most estimators. Encoding and imputation are themselves estimators, which is how they fit into the same interface.
The rule about fit that everything else depends on
fit is where learning happens, and learning must only ever happen on training data. This applies to preprocessing exactly as much as to models.
1from sklearn.preprocessing import StandardScaler23scaler = StandardScaler()4X_train_scaled = scaler.fit_transform(X_train) # LEARN mean and std, then apply5X_test_scaled = scaler.transform(X_test) # APPLY the training values onlyCalling fit_transform on the test set is the error. It looks harmless — the code runs, the shapes are right, the numbers are scaled — but it computes the test set's own mean and standard deviation, which means information about the test data has entered the process. Your reported score improves and your real-world performance does not, which is the worst possible combination because it hides itself.
fit_transformon training data.transformon everything else. If you ever typefitand a variable withtestin its name on the same line, stop.
The mechanism that enforces this automatically is the Pipeline, which chains preprocessing and model into a single object with the same four methods:
1from sklearn.pipeline import Pipeline23pipe = Pipeline([4 ("scale", StandardScaler()),5 ("model", LogisticRegression(max_iter=1000)),6])78pipe.fit(X_train, y_train) # fits the scaler, then the model9print(pipe.score(X_test, y_test)) # transforms with training stats, then predictsBecause the pipeline is itself an estimator, it can be cross-validated or hyperparameter-searched as one unit, and the scaler is refitted inside every fold — which is the only correct way to do it. Doing this by hand is possible and nobody does it correctly for long.
A map of the library
| Module | Holds | Reach for it when |
|---|---|---|
sklearn.datasets | Toy and generated datasets | Testing an idea without real data |
sklearn.model_selection | train_test_split, cross-validation, grid search | Splitting, validating, tuning |
sklearn.preprocessing | Scalers, encoders, binning | Getting features into shape |
sklearn.impute | SimpleImputer, KNNImputer | Filling gaps without leaking |
sklearn.linear_model | Linear and logistic regression, ridge, lasso | A fast, interpretable baseline |
sklearn.tree, sklearn.ensemble | Trees, forests, gradient boosting | Tabular data — usually the strongest option |
sklearn.cluster | K-means, DBSCAN, hierarchical | Finding groups with no labels |
sklearn.decomposition | PCA and friends | Reducing dimensions |
sklearn.metrics | Every scoring function | Judging a model properly |
sklearn.pipeline | Pipeline, ColumnTransformer | Always, in real work |
Naming conventions that let you guess
The library is consistent enough that you can often predict the name you need without looking it up.
| Convention | Means | Example |
|---|---|---|
Suffix Classifier | Predicts categories | RandomForestClassifier |
Suffix Regressor | Predicts numbers | RandomForestRegressor |
| Trailing underscore | An attribute learned during fit | model.coef_, scaler.mean_ |
| No underscore | A setting you chose | model.n_estimators |
random_state= | Makes randomness reproducible | Set it everywhere, always |
The trailing underscore is a genuinely useful signal. model.coef_ does not exist until you call fit — asking for it raises AttributeError, and calling predict on an unfitted model raises NotFittedError — and its presence tells you at a glance which attributes came from the data rather than from your configuration.
Where scikit-learn is the wrong tool
Knowing the boundary saves you from fighting the library.
| Task | Why not scikit-learn | Use instead |
|---|---|---|
| Images, audio, text generation | No deep-learning architectures, little GPU support | PyTorch, TensorFlow |
| Data larger than memory | Estimators expect an in-memory array | Dask, Spark, or sample it |
| Winning tabular competitions | HistGradientBoostingClassifier is fast and strong, but has fewer options and no GPU training | XGBoost, LightGBM, CatBoost (same interface) |
| Time series forecasting | No native seasonality or trend models | statsmodels, Prophet, sktime |
| Statistical inference — p-values, confidence intervals | Built for prediction, not explanation | statsmodels |
| Causal questions — "will this intervention work?" | Correlational by construction | DoWhy, EconML, an experiment |
The statistical inference row surprises people. LinearRegression gives you coefficients but no standard errors and no p-values, because it was designed to predict the next value rather than to test a hypothesis about the world. If your question is "is this effect statistically significant", you want statsmodels.
Note also that XGBoost and LightGBM deliberately copy the scikit-learn interface. So does much of the wider ecosystem. Learning fit/predict/transform buys you far more than one library.
What this looks like when you actually use it
The interface turns model selection into a loop, and a loop is something you can trust and re-run:
1import pandas as pd2from sklearn.pipeline import Pipeline3from sklearn.preprocessing import StandardScaler4from sklearn.model_selection import cross_val_score5from sklearn.linear_model import LogisticRegression6from sklearn.ensemble import RandomForestClassifier, GradientBoostingClassifier7from sklearn.svm import SVC89candidates = {10 "logistic": LogisticRegression(max_iter=1000),11 "svm": SVC(),12 "forest": RandomForestClassifier(random_state=42),13 "boosting": GradientBoostingClassifier(random_state=42),14}1516rows = []17for name, estimator in candidates.items():18 pipe = Pipeline([("scale", StandardScaler()), ("model", estimator)])19 scores = cross_val_score(pipe, X_train, y_train, cv=5, scoring="f1_macro")20 rows.append({"model": name, "mean": scores.mean(), "std": scores.std()})2122print(pd.DataFrame(rows).sort_values("mean", ascending=False).round(3))23# model mean std24# 3 boosting 0.967 0.01725# 1 svm 0.966 0.03226# 0 logistic 0.958 0.02727# 2 forest 0.950 0.017Three things about that output are worth noticing. The scaler is inside the pipeline, so it is refitted within each of the five folds and never sees data it is being evaluated on. The standard deviation is reported next to the mean, because a difference of 0.001 between the top two models is far smaller than the 0.02–0.03 variation between folds — those two are tied, and picking the "winner" is picking noise. On a dataset this small, all four are within noise of each other. And the test set has not been touched at all; every number here comes from splitting the training data.
That is the practical shape of the library. You spend your effort on the parts that actually determine the outcome — the features, the split, the metric, the question — and the estimator becomes a variable you can change with one word.