LLMs Deep Dive

Course Content

LLMs Deep Dive

10 sections · 40 lessons

What are eigenvalues and eigenvectors (in dimensionality reduction)?


What you need to know

Text
A v = λ v

A worked 2 × 2 example

Suppose two standardised features of customer messages have covariance matrix:

Text
C = [[3, 1],     [1, 3]]
  • v1 = [1, 1] / sqrt(2): C·v1 = [4, 4] / sqrt(2) = 4·v1, so λ1 = 4.
  • v2 = [1, -1] / sqrt(2): C·v2 = [2, -2] / sqrt(2) = 2·v2, so λ2 = 2.

Total variance = 4 + 2 = 6. The first component explains 4 / 6 ≈ 67%. Keeping only v1 turns 2 numbers per point into 1, losing a third of the variance.

PCA step by step

  1. Centre the data (subtract each feature's mean); standardise if features use different units.
  2. Covariance — compute the d × d covariance matrix.
  3. Eigen-decompose — eigenvectors are orthogonal directions; eigenvalues are variances.
  4. Choose k — keep enough components to reach, say, 90–95% explained variance (λ_i / sum of λ).
  5. Project — multiply the data by the top-k eigenvectors.
Python
import numpy as npC = np.array([[3.0, 1.0], [1.0, 3.0]])vals, vecs = np.linalg.eigh(C)       # for symmetric matrices, ascending orderprint(vals)                          # [2. 4.]print(vals[::-1] / vals.sum())       # explained variance: [0.667 0.333]

Why SVD in practice

SVD of the centred data matrix gives the same components without forming the covariance matrix, which is more stable numerically and faster for tall data. Libraries such as scikit-learn's PCA use SVD.

Where eigen-ideas appear with LLMs

  • Compressing embeddings — reduce 1,024-dimension vectors to 256 for cheaper storage and search.
  • Visualising — project embeddings to 2D to see clusters (often with UMAP or t-SNE after PCA).
  • Low rank — LoRA assumes fine-tuning updates have most of their energy in a few directions (few large singular values).
  • Stability — the largest singular value of a weight matrix (its spectral norm) bounds how much it can stretch a signal, linking to exploding gradients.
  • Curvature — eigenvalues of the loss Hessian describe how sharp the loss surface is.

A real-life example

An e-commerce search assistant stores a 1,024-dimension embedding for each of 10 million products at 4 bytes per number:

Text
10,000,000 x 1,024 x 4 bytes ≈ 41 GB

The team runs PCA on a sample of 200,000 embeddings. The top 256 components explain about 93% of the variance in their data (illustrative). Projecting to 256 dimensions cuts the index to about 10 GB. Before switching, they measure search quality on 2,000 real queries: recall of the correct product in the top 10 drops by less than a point, which they accept. (Many modern embedding models are trained so you can simply truncate the vector, which is an alternative to PCA.)

Follow-up questions to expect

  • "Why are the principal components orthogonal?" — The covariance matrix is symmetric, and symmetric matrices have orthogonal eigenvectors.
  • "What does a zero eigenvalue mean in PCA?" — No variance in that direction: a feature is a perfect combination of others, so it can be dropped with no loss.
  • "What is the difference between eigenvalues and singular values?" — Singular values exist for any matrix; for a covariance matrix, eigenvalues equal the squared singular values of the centred data divided by n − 1.