Course Content
Machine Learning Essentials
6 sections · 16 lessons
PCA & Dimensionality Reduction
A factory logs 200 sensor readings from one machine every second. Bearing temperature at four points on the housing. Motor current, motor voltage, motor power. Vibration on three axes at five positions. Inlet pressure, outlet pressure, differential pressure.
Look at that list again. Motor power is current times voltage — one column is arithmetic on two others. The four bearing temperatures rise and fall together, because they are four thermometers on the same lump of metal. Differential pressure is the difference of the other two. The dataset has 200 columns and nothing like 200 independent pieces of information.
This causes real problems. You cannot plot 200 dimensions, so you cannot see whether failures cluster. Models trained on highly correlated features produce unstable coefficients. Distance-based methods degrade badly as dimensions multiply. And you are storing and processing far more than you need.
The obvious fix is to drop columns. But which? Drop motor power and you lose a signal that might be cleaner than current and voltage separately. Drop three of the four bearing temperatures and you lose the averaging that was cancelling out sensor noise. Each column carries something, even if most of what it carries duplicates its neighbours.
What you actually want is not to delete columns but to combine them — to replace four correlated temperature readings with one number meaning "how hot the bearing assembly is", constructed so that it keeps as much of the original variation as possible.
The idea: choose better axes
Consider just two features: height in centimetres and weight in kilograms, across 500 people. Plot them and the cloud is an elongated ellipse running diagonally upward — taller people tend to be heavier.
weight weight | . :'. | . :'. | .:'::' | .:'::' <- PC1 runs along here | .:'::' | .:'::' | :' | :' PC2 is perpendicular, +------------- height +------------- height and shortThe original axes are height and weight, and the data varies substantially along both. But look at the ellipse's own shape: it is long in one diagonal direction and thin in the perpendicular one. If you rotated your axes to line up with the ellipse, one new axis would capture almost all the variation and the other almost none.
That first axis is not height and not weight. It is roughly "overall body size" — a mixture of both. The second is roughly "heavy for your height". Those two new features describe the same 500 people using the same two numbers each, but now the first number carries 96% of the information and the second only 4%. Drop the second and you have halved your data while losing almost nothing.
That is principal component analysis: find the directions along which the data varies most, and use those as the new axes.
PCA does not select features. It builds new ones — weighted combinations of all the originals — ordered so that the first captures the most variation, the second the most of what remains, and so on.
How it finds those directions
Step 1: standardise
PCA looks for directions of maximum variance, and variance depends on units. If salary is in pounds (range 20,000–200,000) and years of experience is in years (range 0–40), then salary's variance is astronomically larger and the first principal component will be "salary, essentially". Measure salary in thousands instead and the answer changes completely.
So every feature is centred to mean 0 and scaled to standard deviation 1 first. This is not a tidiness step — PCA on unscaled data is a measurement of your unit choices, not of your data.
Step 2: the covariance matrix
With standardised data X (n rows, p features), the covariance matrix is:
a p×p matrix where entry (i,j) is the covariance between features i and j. On standardised data the diagonal is all 1s and the off-diagonals are correlations.
Step 3: eigenvectors and eigenvalues
An eigenvector of C is a direction that the matrix stretches without rotating:
The eigenvector v is a direction in feature space; the eigenvalue λ is how much variance the data has along it. Larger eigenvalue, more spread. Because C is symmetric, its eigenvectors are mutually perpendicular — the new axes are guaranteed to be at right angles, which means the new features are uncorrelated with each other.
The whole thing, on numbers
Two standardised features correlated at 0.9. The covariance matrix is:
For a matrix of this form the eigenvalues are 1+0.9=1.9 and 1−0.9=0.1, with eigenvectors:
Read those. The first component weights both features equally and positively — it is "both features high together", the overall-size direction. The second weights them oppositely — it is "one high, the other low", the difference direction.
Total variance is 1.9+0.1=2.0, which equals the number of standardised features, as it always does. So:
- PC1 explains 1.9/2.0=95% of the variance.
- PC2 explains 0.1/2.0=5%.
Keep PC1 alone and you have compressed two features into one while retaining 95% of the variation. That is the entire method, and everything at 200 dimensions is this repeated.
Step 4: project
Take the top d eigenvectors as columns of a matrix W and transform:
Each row's new coordinates are its projections onto the chosen directions. In the example, a person standardised to (1.2, 1.4) becomes 21(1.2+1.4)=1.84 on PC1 — a single number replacing two.
Deciding how many components to keep
Three approaches, in rough order of usefulness.
Cumulative explained variance. Pick a threshold — 95% is conventional — and keep however many components reach it. This is the most defensible because the threshold means something.
The scree plot. Plot each eigenvalue in order and look for the elbow where the line flattens. Components after the bend are contributing noise-level amounts.
Downstream performance. If PCA is preprocessing for a model, treat the number of components as a hyperparameter and cross-validate it. This is the only method that optimises what you actually care about.
| Components | Individual variance | Cumulative |
|---|---|---|
| PC1 | 34.2% | 34.2% |
| PC2 | 21.8% | 56.0% |
| PC3 | 14.1% | 70.1% |
| PC4 | 9.3% | 79.4% |
| PC5 | 6.7% | 86.1% |
| PC6 | 4.2% | 90.3% |
| PC7–PC12 | ~1% each | 95.8% |
| PC13–PC200 | < 0.1% each | 100% |
Twelve components out of 200 retain 95.8% of the variance. That is a 94% reduction in width, and it tells you something important about the machine: those 200 sensors are measuring about a dozen underlying physical quantities.
1import numpy as np2from sklearn.decomposition import PCA3from sklearn.preprocessing import StandardScaler45X_scaled = StandardScaler().fit_transform(X_train)67pca = PCA().fit(X_scaled) # all components, to inspect8cum = np.cumsum(pca.explained_variance_ratio_)9print("components for 95%:", np.argmax(cum >= 0.95) + 1)1011# or let scikit-learn choose directly12pca95 = PCA(n_components=0.95).fit(X_scaled)13print("kept:", pca95.n_components_)Passing a float between 0 and 1 to n_components means "keep enough components to reach this proportion of variance", which is usually what you want.
Reading what a component means
Components are combinations of original features, and the weights — called loadings — tell you which features drive each one.
1import pandas as pd23loadings = pd.DataFrame(4 pca.components_[:3].T,5 columns=["PC1", "PC2", "PC3"],6 index=X_train.columns,7)8print(loadings.round(2).sort_values("PC1", key=abs, ascending=False).head(10))A typical result on the sensor data:
| Feature | PC1 | PC2 | PC3 |
|---|---|---|---|
| bearing_temp_1 | 0.31 | −0.04 | 0.12 |
| bearing_temp_2 | 0.30 | −0.02 | 0.09 |
| bearing_temp_3 | 0.29 | −0.06 | 0.14 |
| motor_current | 0.28 | 0.11 | −0.33 |
| vibration_x | 0.04 | 0.42 | 0.08 |
| vibration_y | 0.02 | 0.44 | 0.11 |
| inlet_pressure | 0.06 | −0.09 | 0.51 |
PC1 loads heavily and positively on all the thermal and electrical features — it is "how hard the machine is working". PC2 loads on vibration and almost nothing else — "how much it is shaking". PC3 separates pressure from current — perhaps "flow restriction". Those are interpretable and worth naming.
Two cautions. The sign is arbitrary: eigenvectors are directions, and v and −v describe the same axis, so a component may flip sign between runs or library versions. Only relative signs within a component are meaningful. And loadings sit on standardised features, so they describe influence per standard deviation, not per raw unit.
Using PCA before a model
Two rules, both of which get broken constantly.
Fit PCA on training data only. The components are learned from the covariance structure of whatever data you fit on. Fit on the full dataset and the test set has contributed to defining the axes — leakage, subtle but real.
Scale inside the pipeline. Not before splitting.
1from sklearn.pipeline import make_pipeline2from sklearn.linear_model import LogisticRegression3from sklearn.model_selection import GridSearchCV45pipe = make_pipeline(6 StandardScaler(),7 PCA(n_components=20),8 LogisticRegression(max_iter=1000),9)1011grid = GridSearchCV(12 pipe,13 {"pca__n_components": [5, 10, 20, 40, 80]},14 cv=5, scoring="roc_auc", n_jobs=-1,15).fit(X_train, y_train)1617print(grid.best_params_, round(grid.best_score_, 3))Typical effect on a 200-feature dataset:
| Setup | Features | Fit time | CV AUC |
|---|---|---|---|
| Raw features | 200 | 42 s | 0.883 |
| PCA, 80 components | 80 | 18 s | 0.886 |
| PCA, 20 components | 20 | 5 s | 0.891 |
| PCA, 5 components | 5 | 2 s | 0.842 |
Note that 20 components beat all 200 raw features. That is not magic — the discarded components were mostly measurement noise, and removing them removed something the model was previously fitting. PCA acts as a regulariser here. Push too far, to 5 components, and you start discarding signal.
The caveat that undermines this
PCA is unsupervised. It maximises variance without any knowledge of your target. There is no guarantee that high-variance directions are predictive ones.
A pathological but instructive case: suppose the signal distinguishing your classes lives entirely in a low-variance direction — a small consistent offset — while a large-variance direction reflects irrelevant sensor drift. PCA will keep the drift and discard the signal. This is rare in practice but not vanishingly so, and it is why you should always compare "PCA plus model" against "model on raw features" rather than assuming PCA helps.
What PCA cannot do
| Limitation | Consequence | What to do instead |
|---|---|---|
| Only finds linear combinations | Curved structure (a spiral, a Swiss roll) is not unwrapped | Kernel PCA, UMAP, autoencoders |
| Components are mixtures of everything | "PC3 = 0.51×inlet_pressure − 0.33×current + …" is hard to defend to a regulator | Feature selection, which keeps original columns |
| Ignores the target | May discard the predictive direction | Linear discriminant analysis; or partial least squares |
| Sensitive to outliers | One extreme row can define PC1 by itself | Clip or remove outliers first; robust PCA |
| Assumes variance means information | A noisy sensor gets high variance and dominates | Clean the data; consider ICA if you want independent sources |
| Requires numeric, scaled input | Categorical features need encoding first | MCA for categorical data; or one-hot then PCA, with care |
The interpretability cost is the one most often underestimated. A model on twelve principal components may perform beautifully and be impossible to explain to anyone who needs to sign it off, because no component corresponds to anything a person can point at. When explanation is a requirement, ordinary feature selection — keeping a subset of original columns — is often the better trade even at some cost in accuracy.
When to reach for something else
| Goal | Method | Why |
|---|---|---|
| Preprocessing before a model | PCA | Fast, deterministic, and transforms new data with the same rules |
| A 2-D picture for humans | UMAP or t-SNE | Preserves local neighbourhoods far better in two dimensions |
| Keep original columns | Feature selection | Interpretability survives |
| Nonlinear compression at scale | Autoencoder | Learns arbitrary nonlinear structure |
| Separate mixed sources | ICA | Seeks statistical independence, not maximum variance |
A warning about t-SNE and UMAP, since they produce the prettiest pictures: distances between clusters in their output are not meaningful, and neither is cluster size. Two clusters far apart in a t-SNE plot are not necessarily more different than two that are close. Use them to see whether structure exists, never to measure it.
Compression, made tangible
PCA on an image dataset makes the trade-off visible. Treat each image as a row of pixel values and reduce.
1from sklearn.datasets import load_digits2from sklearn.decomposition import PCA3import numpy as np45digits = load_digits()6X = digits.data # 1797 images, 64 pixels each78for k in [2, 5, 10, 20, 40, 64]:9 pca = PCA(n_components=k).fit(X)10 reconstructed = pca.inverse_transform(pca.transform(X))11 err = np.mean((X - reconstructed) ** 2)12 print(f"{k:2d} components: {pca.explained_variance_ratio_.sum():.1%} "13 f"variance, reconstruction MSE {err:6.2f}, "14 f"storage {k / 64:.0%}")inverse_transform is what makes this concrete: it maps the reduced representation back into the original 64-dimensional space. What comes back is not the original — the discarded components are gone forever — but at 20 components the digits are clearly readable while occupying 31% of the storage. The information you lost was fine detail and sensor noise; the information you kept was shape.
What this means when you build something
Reach for PCA when width is causing a concrete problem: a model that is too slow, features too correlated for stable coefficients, a distance-based method degrading, or a dataset you cannot visualise. If none of those apply, adding PCA adds a step and removes interpretability for nothing.
When you do use it, put StandardScaler immediately before it in a pipeline, without exception, and treat the number of components as a hyperparameter to cross-validate rather than a number to pick by staring at a scree plot. Then run the model on raw features as well and compare honestly — PCA earns its place by beating the alternative, not by being applied.
And spend ten minutes on the loadings even if you never show them to anyone. Discovering that 200 sensor columns collapse to twelve components, and that the first is "working hard" and the second is "shaking", tells you something true about the machine that no accuracy figure will. That understanding usually suggests better features than PCA itself produced.