Course Content
Data Science Fundamentals
4 sections · 10 lessons
Dimensionality Reduction with PCA
A customer satisfaction survey asks 40 questions. "Was the site easy to navigate?" "Could you find what you needed?" "Was the layout clear?" "Did the search work well?" Four separate columns — and if you look at their correlation matrix, they sit between 0.82 and 0.91 with each other. They are not four measurements. They are one measurement, asked four ways.
Across the whole survey the same thing happens repeatedly. Perhaps 40 columns are really tracking five underlying attitudes: ease of use, price satisfaction, trust, support quality, delivery experience. Your data has 40 dimensions and roughly five dimensions' worth of information.
You could pick one question from each cluster and throw the other 35 away. That works, and it wastes the extra information in the discarded questions — each of the four navigation questions catches a slightly different aspect, and averaging them cancels some of the random noise in individual responses. What you want is a small set of new columns, each combining the originals, that keeps as much of the real variation as possible. That is what principal component analysis does.
The Idea: Find the Direction That Matters Most
Forget 40 dimensions and picture two: height and weight, measured across a thousand people. Plotted, the points form an elongated cloud running diagonally upwards, because taller people generally weigh more.
weight ^ . ' | . ' : | . ' : . <- most variation runs along here | . ' : . | . ' : . <- very little runs across here +------------------------------> heightThe cloud is long in one direction and thin in the other. If you had to describe each person with a single number instead of two, the useful number is where they sit along the long axis — roughly, overall body size. The position across the thin axis says comparatively little, because everyone has almost the same value for it.
That long axis is the first principal component. It is not height and it is not weight; it is a specific mixture, perhaps 0.71 × height + 0.70 × weight after both are standardised. The second component is the direction perpendicular to it that captures the most of what is left — here, something like "heavy for your height".
PCA does not choose which of your columns to keep. It builds new columns, each a weighted blend of all the originals, ordered by how much of the data's variation each one captures.
The Recipe
- Standardise every column to mean 0 and standard deviation 1.
- Compute the covariance matrix of the standardised data.
- Find its eigenvectors and eigenvalues.
- Sort by eigenvalue, largest first. Each eigenvector is a component direction; its eigenvalue is the variance along that direction.
- Keep the first k, and project the data onto them.
The covariance between two standardised variables is just their correlation, so step two is asking "which variables move together, and how strongly?". The eigenvectors of that matrix are the directions along which the data varies independently — mutually perpendicular, so the new components are uncorrelated with each other by construction. The eigenvalue attached to each tells you how much variance lives in that direction:
Why Standardising First Is Not Optional
PCA maximises variance, and variance depends on units. Consider three columns: annual income in pounds (standard deviation 25,000), age in years (standard deviation 14), and a satisfaction score from 1 to 5 (standard deviation 1.2).
Variance contributed, unstandardised: income 25,000^2 = 625,000,000 99.99997% age 14^2 = 196 0.00003% satisfaction 1.2^2 = 1.44 0.0000002%The first principal component will be income, essentially unmixed, because nothing else registers on that scale. Change income to thousands of pounds and the entire analysis changes — same data, different units, different answer. That instability is the tell that the procedure is measuring your unit choices rather than your data.
1import numpy as np2from sklearn.preprocessing import StandardScaler3from sklearn.decomposition import PCA45X_std = StandardScaler().fit_transform(X) # do this first, every time6pca = PCA().fit(X_std)After standardising, every column contributes variance 1, so all 40 survey questions start equal and the components reflect genuine correlation structure. The rare exception is data already in identical, meaningful units — pixel intensities, say, or a set of prices in the same currency — where relative magnitude is real information you may want to preserve.
Doing It by Hand
Working through the linear algebra once makes the rest concrete.
1import numpy as np23rng = np.random.default_rng(42)4n = 5005height = rng.normal(170, 10, n)6weight = 0.9 * height - 90 + rng.normal(0, 5, n)7X = np.column_stack([height, weight])89# 1. Centre and scale10X_std = (X - X.mean(axis=0)) / X.std(axis=0, ddof=1)1112# 2. Covariance matrix (2x2 here)13cov = np.cov(X_std, rowvar=False)14print(cov.round(3))15# [[1. 0.86]16# [0.86 1. ]]1718# 3. Eigendecomposition (eigh is for symmetric matrices)19eigvals, eigvecs = np.linalg.eigh(cov)2021# 4. Sort descending22order = np.argsort(eigvals)[::-1]23eigvals, eigvecs = eigvals[order], eigvecs[:, order]24print("eigenvalues:", eigvals.round(3)) # [1.86 0.14]25print("explained:", (eigvals / eigvals.sum()).round(4)) # [0.9302 0.0698]2627# 5. Project28X_pca = X_std @ eigvecs[:, :1] # keep one componentFor a 2×2 correlation matrix with correlation r, the eigenvalues are exactly 1 + r and 1 − r. This sample has r = 0.860, which gives 1.860 and 0.140, matching the output. The first component captures 93.0% of the variation in two columns using one number. The corresponding eigenvector is approximately [0.707, 0.707] — an equal blend, which is what "overall size" should look like when both variables are standardised.
Two things to know about the output. Component signs are arbitrary: an eigenvector multiplied by −1 is an equally valid eigenvector, so a component may flip direction between runs or library versions without anything being wrong. And scikit-learn uses singular value decomposition rather than explicit eigendecomposition — numerically more stable, mathematically equivalent.
How Many Components to Keep
By variance target
1pca_full = PCA().fit(X_std)2cum = np.cumsum(pca_full.explained_variance_ratio_)34for target in (0.80, 0.90, 0.95, 0.99):5 print(f"{target:.0%} variance -> {np.searchsorted(cum, target) + 1} components")67# Or let scikit-learn choose8pca = PCA(n_components=0.95, svd_solver="full").fit(X_std)9print("kept:", pca.n_components_)Passing a float between 0 and 1 as n_components means "keep enough components to reach this much variance". 95% is a common default; 99% for a compression task where fidelity matters; 80% is often enough when the components feed a model that will do its own work.
By the elbow in a scree plot
1import matplotlib.pyplot as plt23fig, ax = plt.subplots(1, 2, figsize=(13, 4.5))4ax[0].plot(range(1, len(pca_full.explained_variance_ratio_) + 1),5 pca_full.explained_variance_ratio_, "o-")6ax[0].set(xlabel="Component", ylabel="Explained variance ratio", title="Scree plot")78ax[1].plot(range(1, len(cum) + 1), cum, "o-")9ax[1].axhline(0.95, ls="--", color="crimson")10ax[1].set(xlabel="Components", ylabel="Cumulative variance")The scree plot usually drops steeply then flattens. The point where it levels off is the elbow: components past it each add a sliver of variance, which is often mostly noise. On the 40-question survey you would expect a clear elbow around five.
| Rule | Keep | Best for |
|---|---|---|
| Variance threshold | Enough for 90–95% | General-purpose compression |
| Scree elbow | Up to where the curve flattens | Finding latent structure |
| Kaiser criterion | Eigenvalue > 1 (on standardised data) | Survey and psychometric work |
| Cross-validation | Whatever maximises held-out score | PCA as a modelling step |
| Fixed at 2 or 3 | — | Plotting only |
The Kaiser rule has a neat justification: on standardised data every original variable has variance 1, so a component with an eigenvalue below 1 explains less than a single original column would. It is a crude rule and it tends to keep too many, but it is a reasonable floor.
Reading the Components
A component is only interpretable through its loadings — the weight each original variable receives.
1import pandas as pd23loadings = pd.DataFrame(4 pca.components_.T,5 index=feature_names,6 columns=[f"PC{i+1}" for i in range(pca.n_components_)],7)89for pc in loadings.columns[:3]:10 top = loadings[pc].abs().sort_values(ascending=False).head(6).index11 print(f"\n{pc} ({pca.explained_variance_ratio_[int(pc[2:])-1]:.1%} of variance)")12 print(loadings.loc[top, pc].round(3))On the survey data you might see PC1 loading 0.28 to 0.34 on every navigation and search question and near zero elsewhere — that component is "ease of use". PC2 might load +0.4 on price questions and −0.35 on premium-feature questions, making it a value-for-money axis where positive means price-sensitive.
Read the signs together. A component with positive loadings on some variables and negative on others is a contrast: it measures the difference between two groups of variables, not the level of either. Reading only the magnitudes will get this exactly backwards for half the respondents.
A principal component with no plausible interpretation is still mathematically valid, but you have lost the ability to explain your model to anyone. That is the price PCA charges.
PCA for Visualisation
Projecting to two components is the fastest way to see whether a high-dimensional dataset has structure at all.
1p2 = PCA(n_components=2, random_state=42)2coords = p2.fit_transform(X_std)34plt.figure(figsize=(8, 6))5for label in np.unique(y):6 m = y == label7 plt.scatter(coords[m, 0], coords[m, 1], label=str(label), alpha=0.6, s=18)8plt.xlabel(f"PC1 ({p2.explained_variance_ratio_[0]:.1%})")9plt.ylabel(f"PC2 ({p2.explained_variance_ratio_[1]:.1%})")10plt.legend()Always put the explained-variance percentages on the axis labels. If PC1 and PC2 together account for 78%, the picture is broadly trustworthy. If they account for 21%, you are looking at a shadow — groups that appear to overlap may be cleanly separated in the 79% of variance you cannot see, and concluding "the classes are not separable" from such a plot is a genuine mistake.
For visualisation specifically, methods like t-SNE and UMAP usually reveal cluster structure better, because they are built to preserve local neighbourhoods rather than global variance. They are also slower and their output cannot be interpreted as directions in the original space. A common pattern is PCA down to 50 components first, then UMAP on those — fast, and it removes noise before the expensive step.
Using PCA Inside a Model
PCA learns parameters — the means, the scales, the component directions — so it must be fitted on training data only. Fitted on everything, it has seen the test set's structure, and your validation score becomes optimistic.
1from sklearn.pipeline import Pipeline2from sklearn.model_selection import cross_val_score, GridSearchCV3from sklearn.linear_model import LogisticRegression45pipe = Pipeline([6 ("scale", StandardScaler()),7 ("pca", PCA(random_state=42)),8 ("clf", LogisticRegression(max_iter=1000)),9])1011grid = GridSearchCV(pipe, {"pca__n_components": [5, 10, 20, 40]},12 cv=5, scoring="roc_auc")13grid.fit(X_train, y_train)14print(grid.best_params_, round(grid.best_score_, 4))Tuning n_components by cross-validation is more honest than picking a variance threshold, because it optimises the thing you actually care about. Note that the components maximising variance are not necessarily the ones that predict your target — variance is computed without ever looking at y. It is entirely possible for the signal you need to sit in component 12, and dropping everything after component 10 discards it. This is the most important limitation of PCA as a feature-selection step.
For data too large to hold in memory:
1from sklearn.decomposition import IncrementalPCA23ipca = IncrementalPCA(n_components=50, batch_size=1000)4for chunk in pd.read_csv("huge.csv", chunksize=1000):5 ipca.partial_fit(scaler.transform(chunk[feature_names]))When Not to Use It
| Situation | PCA? | Reasoning |
|---|---|---|
| Many correlated numeric features | Yes | Exactly what it is for |
| Speeding up a distance-based model | Yes | Fewer dimensions, less noise |
| Visualising high-dimensional data | Yes | Fast first look |
| Features are already uncorrelated | No | Nothing to compress; components ≈ original columns |
| You must explain the model | No | "Component 3 went up" satisfies no regulator |
| Mostly categorical data | No | Variance on one-hot columns is not meaningful; use MCA or embeddings |
| Relationships are strongly non-linear | Careful | PCA only finds linear structure; consider kernel PCA or autoencoders |
| Using tree-based models | Usually no | Trees handle irrelevant features well, and rotation destroys axis-aligned splits |
That last row deserves a note, because it runs against habit. Decision trees split on one feature at a time along the original axes. If a clean rule exists — "approve when income > 40,000" — a tree finds it instantly. After PCA, income is smeared across every component, and the tree needs many oblique-ish splits to approximate a rule it could previously express in one. Applying PCA before a random forest frequently makes performance worse.
What This Means When You Build Something
The mistake that costs the most is quiet: fitting the scaler and PCA on the full dataset, then splitting. It never errors and it inflates every score you report. Keep both steps inside a pipeline and the problem cannot occur.
Then check that the reduction was worth it, rather than assuming:
1from sklearn.model_selection import cross_val_score23baseline = Pipeline([("scale", StandardScaler()),4 ("clf", LogisticRegression(max_iter=1000))])5reduced = Pipeline([("scale", StandardScaler()),6 ("pca", PCA(n_components=0.95, random_state=42)),7 ("clf", LogisticRegression(max_iter=1000))])89for name, p in [("all features", baseline), ("PCA 95%", reduced)]:10 s = cross_val_score(p, X_train, y_train, cv=5, scoring="roc_auc")11 print(f"{name:14s} {s.mean():.4f} +/- {s.std():.4f}")Run this before committing. PCA earns its place when it makes the model faster, more stable, or better on held-out data — and when it does not, you have traded away every ounce of interpretability for nothing. On the 40-question survey it earns its place easily: five components carrying 87% of the variation, each one a nameable attitude, replacing 40 columns that were mostly repeating each other. That is the case where PCA is not a technique applied to data but a discovery about it.