Mathematics for Machine Learning

Determinants, Matrix Inverse, and Solving Linear Systems


Here is a system of two equations in two unknowns. It looks like every textbook example you have ever seen.

2x+3y=82x + 3y = 8

4x+6y=164x + 6y = 16

Two equations, two unknowns, so there should be exactly one answer. Try to find it. Double the first equation and you get 4x+6y=164x + 6y = 16 — which is the second equation, character for character. The second row told you nothing the first row had not already said. So x=1,y=2x = 1, y = 2 works. So does x=4,y=0x = 4, y = 0. So does x=−2.5,y=13/3x = -2.5, y = 13/3. There are infinitely many solutions and no way to choose between them.

Now change one digit. Make the right-hand side of the second equation 17 instead of 16. The equations now claim that the same quantity 4x+6y4x + 6y equals both 16 and 17. There are zero solutions. The same left-hand side gave us infinitely many answers and then none, depending on a number that never appeared in the coefficients at all.

You could spot this by eye with two equations. With a 500-variable system inside a regression fit, you cannot. You need a single number, computed from the coefficients alone, that tells you in advance whether a unique solution exists. That number is the determinant, and understanding what it actually measures explains not only when systems break but why some models are numerically fragile long before they break outright.

The determinant is a warning light, not a verdictDeterminant well clear of zero• Columns span the space independently• The transformation can be undone• One unique solution exists• Small input errors stay smallDeterminant at or near zero• Columns are collinear or dependent• Area collapses, information is lost• No solution, or infinitely many• Near-zero is worse: it answers, wrongly
A singular matrix fails loudly; an ill-conditioned one returns confident numbers that move wildly with the data.

A matrix is a transformation of space

Stop reading AxA\mathbf{x} as "a table times a list" and start reading it as "a machine that moves points". Feed in a vector, get out a different vector. Feed in every point in the plane and the whole plane gets reshaped in a specific, rigid way: straight lines stay straight, the origin stays put, and evenly spaced points stay evenly spaced. That is what makes it a linear transformation.

  • [3002]\begin{bmatrix} 3 & 0 \\ 0 & 2 \end{bmatrix} is a scaling: it stretches everything three times horizontally and twice vertically. Reversible — scale back by 1/31/3 and 1/21/2.
  • [0−110]\begin{bmatrix} 0 & -1 \\ 1 & 0 \end{bmatrix} is a rotation by 90 degrees anticlockwise. Reversible — rotate the other way.
  • [10.501]\begin{bmatrix} 1 & 0.5 \\ 0 & 1 \end{bmatrix} is a shear: it slides points sideways in proportion to their height, turning a square into a leaning parallelogram. Reversible — shear back.
  • [1224]\begin{bmatrix} 1 & 2 \\ 2 & 4 \end{bmatrix} is none of those. It squashes the entire plane onto a single line, and it is not reversible.

The last one is the interesting case. Every input point lands somewhere on the line y=2xy = 2x. The points (1,0)(1,0) and (0,0.5)(0,0.5) both map to (1,2)(1,2). Once two different inputs produce the same output, no rule on earth can undo the transformation, because the output does not say which input it came from. Information has been destroyed.

That is the same failure as the two-equations problem at the top, wearing different clothes. Solving Ax=bA\mathbf{x} = \mathbf{b} means asking "which point lands on b\mathbf{b}?" If the transformation flattened space onto a line, then points on that line came from infinitely many places, and points off it came from nowhere at all.

The determinant measures how much space is squashed

Take the unit square — corners at (0,0),(1,0),(0,1),(1,1)(0,0), (1,0), (0,1), (1,1), area exactly 1. Push it through the transformation. It comes out as a parallelogram. The determinant is the area of that parallelogram.

So det⁡A=3\det A = 3 means every region's area triples. det⁡A=0.5\det A = 0.5 means areas halve. A negative determinant means space was flipped over, like turning a page, and the magnitude still gives the area factor. In three dimensions, swap "area" for "volume"; in nn dimensions the same idea holds with a word nobody needs.

The determinant is the factor by which a transformation multiplies area. A determinant of zero means area was multiplied by zero — space collapsed, and nothing collapsed can be un-collapsed.

For a 2×22 \times 2 matrix the formula is short:

det⁡[abcd]=ad−bc\det \begin{bmatrix} a & b \\ c & d \end{bmatrix} = ad - bc

With A=[3124]A = \begin{bmatrix} 3 & 1 \\ 2 & 4 \end{bmatrix}: det⁡A=(3)(4)−(1)(2)=12−2=10\det A = (3)(4) - (1)(2) = 12 - 2 = 10. Areas grow tenfold. Now the coefficient matrix from the opening system, [2346]\begin{bmatrix} 2 & 3 \\ 4 & 6 \end{bmatrix}: (2)(6)−(3)(4)=12−12=0(2)(6) - (3)(4) = 12 - 12 = 0. There it is. The determinant knew the system was degenerate without anyone attempting to solve it.

Three by three, by cofactor expansion

For larger matrices you expand along a row: take each entry, multiply it by the determinant of the smaller matrix left when you delete that entry's row and column, and alternate the signs +,−,++, -, +.

A=[6114−25287]A = \begin{bmatrix} 6 & 1 & 1 \\ 4 & -2 & 5 \\ 2 & 8 & 7 \end{bmatrix}

Expanding along the top row:

det⁡A=6∣−2587∣−1∣4527∣+1∣4−228∣\det A = 6\begin{vmatrix} -2 & 5 \\ 8 & 7 \end{vmatrix} - 1\begin{vmatrix} 4 & 5 \\ 2 & 7 \end{vmatrix} + 1\begin{vmatrix} 4 & -2 \\ 2 & 8 \end{vmatrix}

Each small determinant is ad−bcad - bc. First: (−2)(7)−(5)(8)=−14−40=−54(-2)(7) - (5)(8) = -14 - 40 = -54. Second: (4)(7)−(5)(2)=28−10=18(4)(7) - (5)(2) = 28 - 10 = 18. Third: (4)(8)−(−2)(2)=32+4=36(4)(8) - (-2)(2) = 32 + 4 = 36. Putting it together:

det⁡A=6(−54)−1(18)+1(36)=−324−18+36=−306\det A = 6(-54) - 1(18) + 1(36) = -324 - 18 + 36 = -306

Volumes are multiplied by 306 and space is flipped inside out. Note the cost, though: a 3×33\times3 needs three 2×22\times2s, a 4×44\times4 needs four 3×33\times3s, and this grows like n!n!. Nobody computes large determinants this way; software reduces the matrix to triangular form first, where the determinant is just the product of the diagonal.

Properties worth memorising

PropertyWhy it is true, in one line
det⁡(AB)=det⁡A⋅det⁡B\det(AB) = \det A \cdot \det BApply one transformation then another; the area factors multiply
det⁡(AT)=det⁡A\det(A^T) = \det ATransposing does not change how much space the map scales
Triangular matrix: det⁡=∏iAii\det = \prod_i A_{ii}Each diagonal entry stretches one axis; the off-diagonal shear costs no area
det⁡A=0⇒A\det A = 0 \Rightarrow A is singularSpace collapsed to lower dimension; no inverse exists
det⁡(A−1)=1/det⁡A\det(A^{-1}) = 1/\det AUndoing an area tripling means dividing area by three

Singular is simply the technical word for "flattens space, cannot be undone". When a library raises LinAlgError: Singular matrix, this is what it found: two of your features are exact linear combinations of each other, or you have fewer independent samples than parameters, and the maths has nowhere to go.

The inverse: the transformation that undoes the transformation

A−1A^{-1} is the matrix that puts every point back where it started:

AA−1=A−1A=IAA^{-1} = A^{-1}A = I

where II is the identity — the do-nothing transformation with 1s down the diagonal. An inverse exists exactly when det⁡A≠0\det A \neq 0, which is the same statement as "no information was destroyed", which is the same statement as "Ax=bA\mathbf{x} = \mathbf{b} has one unique solution for every b\mathbf{b}". Three phrasings, one fact.

For 2×22 \times 2 there is a closed form: swap the diagonal entries, negate the off-diagonal ones, divide by the determinant.

A−1=1ad−bc[d−b−ca]A^{-1} = \frac{1}{ad - bc}\begin{bmatrix} d & -b \\ -c & a \end{bmatrix}

Take A=[4726]A = \begin{bmatrix} 4 & 7 \\ 2 & 6 \end{bmatrix}. The determinant is (4)(6)−(7)(2)=24−14=10(4)(6) - (7)(2) = 24 - 14 = 10, so

A−1=110[6−7−24]=[0.6−0.7−0.20.4]A^{-1} = \frac{1}{10}\begin{bmatrix} 6 & -7 \\ -2 & 4 \end{bmatrix} = \begin{bmatrix} 0.6 & -0.7 \\ -0.2 & 0.4 \end{bmatrix}

Verify by multiplying. Top-left: 4(0.6)+7(−0.2)=2.4−1.4=14(0.6) + 7(-0.2) = 2.4 - 1.4 = 1. Top-right: 4(−0.7)+7(0.4)=−2.8+2.8=04(-0.7) + 7(0.4) = -2.8 + 2.8 = 0. Bottom-left: 2(0.6)+6(−0.2)=1.2−1.2=02(0.6) + 6(-0.2) = 1.2 - 1.2 = 0. Bottom-right: 2(−0.7)+6(0.4)=−1.4+2.4=12(-0.7) + 6(0.4) = -1.4 + 2.4 = 1. That is II. Notice where the determinant sits: in the denominator. As det⁡A\det A approaches zero the entries of the inverse blow up, and that is where numerical trouble begins.

Beyond 2×22 \times 2 there is no memorable formula. The practical method is Gauss-Jordan elimination: write AA and II side by side, then use row operations — swap two rows, scale a row, add a multiple of one row to another — to grind the left block into II. Whatever the right block has become is A−1A^{-1}, because you applied the same undo-sequence to both.

Solving Ax=bA\mathbf{x} = \mathbf{b} the right way

The algebra says: multiply both sides by A−1A^{-1} and get x=A−1b\mathbf{x} = A^{-1}\mathbf{b}. The algebra is correct and the code is wrong.

Python
import numpy as npx = np.linalg.inv(A) @ b     # correct maths, bad numericsx = np.linalg.solve(A, b)    # what you should actually write

solve never builds the inverse. It performs LU decomposition — factoring AA into a lower-triangular LL times an upper-triangular UU — and then solves two triangular systems by simple substitution. The difference is not cosmetic:

x=A−1b\mathbf{x} = A^{-1}\mathbf{b}np.linalg.solve(A, b)
CostAbout 2n32n^3 operations, then a multiplyAbout 23n3\tfrac{2}{3}n^3 — roughly three times cheaper
Rounding errorAccumulated twice: forming the inverse, then using itAccumulated once, with pivoting to keep it small
Sparse matricesInverse is usually dense; memory explodesFactors stay sparse
Near-singular AASilently returns huge, meaningless numbersAlso silent, but stays as accurate as the problem allows

Seeing A−1A^{-1} in a formula is a licence to solve a system, not an instruction to compute an inverse.

Where this turns up in machine learning

Fitting a linear regression by least squares has an exact closed-form answer, the normal equation:

w=(XTX)−1XTy\mathbf{w} = (X^TX)^{-1}X^T\mathbf{y}

In words: XX is your data with samples in rows and features in columns, y\mathbf{y} the targets, and w\mathbf{w} the weights that minimise total squared error. XTXX^TX is a small d×dd \times d matrix summarising how features co-vary; inverting it undoes that overlap so each feature gets credit only for what it uniquely explains. This is the formula every textbook prints and the one you should never type literally. If two features are perfectly correlated — height in centimetres and height in inches — then XTXX^TX is singular, the inverse does not exist, and the fit has no unique answer. That is exactly the doubled-equation problem from the opening, arriving through a data-cleaning mistake. Use np.linalg.lstsq, or add a ridge penalty, which adds λI\lambda I to XTXX^TX and nudges the determinant away from zero.

The other place is covariance matrices. Centre your data, compute C=1n−1XTXC = \frac{1}{n-1}X^TX, and entry CijC_{ij} says how features ii and jj vary together. Principal component analysis works entirely on this matrix. A zero determinant here means some feature is a perfect combination of others and one direction in your data carries no variance whatsoever — real information, if you know to look for it.

Conditioning: the failure that is worse than failing

Outright singularity at least announces itself. The dangerous case is nearly singular, where everything runs and the answers are rubbish.

A=[1111.001],det⁡A=1.001−1=0.001A = \begin{bmatrix} 1 & 1 \\ 1 & 1.001 \end{bmatrix}, \qquad \det A = 1.001 - 1 = 0.001

Solve Ax=[22.001]A\mathbf{x} = \begin{bmatrix} 2 \\ 2.001 \end{bmatrix}. Subtracting the first equation from the second gives 0.001y=0.0010.001y = 0.001, so y=1y = 1 and x=1x = 1.

Now nudge the second entry of b\mathbf{b} from 2.0012.001 to 2.0022.002 — a change of 0.05%, smaller than any real measurement error. Subtracting gives 0.001y=0.0020.001y = 0.002, so y=2y = 2 and x=0x = 0. The solution moved from (1,1)(1,1) to (0,2)(0,2). A rounding-level wobble in the input produced a 100% change in the output.

The quantity that predicts this is the condition number:

κ(A)=σmax⁡σmin⁡\kappa(A) = \frac{\sigma_{\max}}{\sigma_{\min}}

where σ\sigma are the singular values — the stretch factors of the transformation along its most and least stretched directions. So κ\kappa asks: how lopsided is this map? For our matrix, σmax⁡≈2.0005\sigma_{\max} \approx 2.0005 and σmin⁡≈0.0005\sigma_{\min} \approx 0.0005, giving κ≈4002\kappa \approx 4002. The rule of thumb is that you lose about log⁡10κ\log_{10}\kappa digits of accuracy, so roughly 3.6 digits gone, from the 16 that double precision gives you.

κ(A)\kappa(A)ReadingWhat to do
1 to 10Well conditionedNothing; trust the answer
10310^3 to 10610^6Getting uncomfortableScale features to comparable ranges
Above 101010^{10}Effectively singular in double precisionRegularise, drop redundant features, or use an SVD-based solver

Feature scaling is the cheapest fix and the most neglected. If one column is measured in millions and another in hundredths, κ\kappa is enormous before the model has seen a single data point. Standardising every column to zero mean and unit variance often drops the condition number by several orders of magnitude and costs one line of code.

What to do when you actually build this

Everything above compresses into a short checklist for numerical code that has to work on real data.

Compute np.linalg.cond(A) before you trust any linear solve, and log it. It is one cheap call and it tells you whether the answer you are about to print is meaningful to sixteen digits or to three. Never write np.linalg.inv, because it is slower, less accurate, and almost always a sign that you have transcribed a formula instead of thinking about it — reach for solve, lstsq, or a factorisation. Standardise your features, both because models train better and because it fixes conditioning at the root. And when a solver reports a singular matrix, do not reach for a pseudo-inverse to make the message go away: go and look for the duplicated column, the constant feature, or the one-hot encoding where you forgot to drop a category. The determinant is telling you something true about your data.

The number ad−bcad - bc that looked like an arbitrary bit of arithmetic is, in the end, a smoke alarm for your whole pipeline. Learn to hear it.