Mathematics for Machine Learning

Vectors, Matrices, and Operations


You have a price model for houses. It is simple: each listing has 50 numeric features — floor area, bedroom count, year built and so on — and you multiply each one by its own weight, add a baseline, and out comes a predicted price. You have a million listings to score. So you write the obvious thing.

Python
predictions = []for listing in listings:              # 1,000,000 listings, 50 features each    total = intercept    for feature, weight in zip(listing, weights):        total += feature * weight    predictions.append(total)

It works. It also takes five or six seconds, even on a fast laptop. That is fine once, and painful when you are tuning the model and want to re-score every time you change a weight. Now watch the same computation written differently:

Python
predictions = X @ w + b               # done in about 7 milliseconds

Same arithmetic. Same answers. Several hundred times faster, because that one line hands the whole job to a compiled routine that keeps the CPU's vector units busy and never touches the Python interpreter. That gap is not a trick. It is the reason machine learning is written in the language of vectors and matrices rather than loops. Learn to see your data as arrays of numbers and the operations on them as a handful of named products, and you get speed, clarity, and a vocabulary shared by every ML library ever written. Here is that vocabulary, from the bottom.

Scoring 100,000 listings is one shape calculationX is100,000 by 3w is 3 by 1Innerdimensions3 and 3 agreePrediction is100,000 by 1The loop over listings disappears into the inner dimension that gets summed away.
Batching a model is not an optimisation trick — it is the definition of matrix multiplication.

A vector is a list of numbers that means something

A vector is an ordered list of numbers. That is the whole definition. What makes it useful is that the order carries meaning: slot one is always floor area, slot two is always bedrooms, and every vector in your dataset agrees on that.

x=[140031995]\mathbf{x} = \begin{bmatrix} 1400 \\ 3 \\ 1995 \end{bmatrix}

This is one house: 1400 square feet, 3 bedrooms, built in 1995. Because it has three numbers we say it lives in R3\mathbb{R}^3 — three-dimensional real space. A vector with 768 numbers lives in R768\mathbb{R}^{768}, which you cannot picture and never need to.

There are three ways to read the same object, and switching between them freely is most of the skill:

ReadingWhat it saysWhen it helps
A listThree stored numbers, nothing moreWriting code, thinking about memory and shapes
A pointA location in space, one axis per featureDistance, clustering, nearest neighbours
An arrowA direction plus a magnitude, from the originGradients, similarity, anything about angle

In machine learning the word you will meet is feature vector: one row of your dataset, one example, with every measurement about it laid out in a fixed order. A house, a customer, a sentence turned into 768 numbers by a language model — all the same shape of object.

Adding vectors and scaling them

Two vectors of the same length add position by position:

[34]+[12]=[46]\begin{bmatrix} 3 \\ 4 \end{bmatrix} + \begin{bmatrix} 1 \\ 2 \end{bmatrix} = \begin{bmatrix} 4 \\ 6 \end{bmatrix}

Multiplying by a single number — a scalar — stretches every component by the same factor: 2⋅[3,4]=[6,8]2 \cdot [3, 4] = [6, 8]. The arrow keeps its direction and doubles in length. A negative scalar flips it round. These two operations are the entire content of gradient descent: your weights are a vector, the gradient is a vector, and the update w←w−η∇\mathbf{w} \leftarrow \mathbf{w} - \eta \nabla is a scalar multiply followed by a subtraction.

The dot product, and why it is really about angle

The dot product multiplies matching components and sums the results:

a⋅b=∑i=1naibi\mathbf{a} \cdot \mathbf{b} = \sum_{i=1}^{n} a_i b_i

With a=[3,4]\mathbf{a} = [3, 4] and b=[1,2]\mathbf{b} = [1, 2]: 3×1+4×2=3+8=113 \times 1 + 4 \times 2 = 3 + 8 = 11. One number out, no matter how long the vectors are.

That arithmetic hides something geometric. The same quantity equals

a⋅b=∥a∥ ∥b∥cos⁡θ\mathbf{a} \cdot \mathbf{b} = \|\mathbf{a}\| \, \|\mathbf{b}\| \cos\theta

where ∥⋅∥\|\cdot\| means length and θ\theta is the angle between the two arrows. In plain English: the dot product is big when the vectors are long and point the same way, zero when they are at right angles, and negative when they point apart. Check it on our numbers. ∥a∥=9+16=5\|\mathbf{a}\| = \sqrt{9+16} = 5 and ∥b∥=1+4≈2.236\|\mathbf{b}\| = \sqrt{1+4} \approx 2.236, so cos⁡θ=11/(5×2.236)=0.984\cos\theta = 11 / (5 \times 2.236) = 0.984, giving θ≈10.3∘\theta \approx 10.3^\circ. Nearly the same direction, which matches the picture.

Rearrange that formula and you have cosine similarity, the standard way to ask "are these two things alike?":

cos⁡θ=a⋅b∥a∥ ∥b∥\cos\theta = \frac{\mathbf{a} \cdot \mathbf{b}}{\|\mathbf{a}\|\,\|\mathbf{b}\|}

Dividing by both lengths strips out magnitude and leaves only direction, on a scale from −1-1 to 11. This is why search engines and recommendation systems compare embeddings with cosine similarity rather than raw dot products: a long document should not beat a short one simply for being long.

The dot product is the only bridge between arithmetic and geometry you need: multiply-and-sum on one side, angle on the other.

Length, and why models care about it

The length of a vector — its norm, written ∥x∥\|\mathbf{x}\| — is Pythagoras in as many dimensions as you like:

∥x∥2=∑ixi2\|\mathbf{x}\|_2 = \sqrt{\sum_i x_i^2}

For [3,4][3, 4] that is 9+16=5\sqrt{9 + 16} = 5. This is the L2 norm, the everyday notion of distance. Its sibling the L1 norm adds absolute values instead, ∥[3,4]∥1=3+4=7\|[3,4]\|_1 = 3 + 4 = 7 — the distance you would walk on a street grid. The choice matters in regularisation: penalising the L2 norm of the weights shrinks them all smoothly, while penalising the L1 norm drives many of them to exactly zero and gives you feature selection for free.

To normalise a vector, divide it by its own length:

x^=x∥x∥so[3,4]5=[0.6, 0.8]\hat{\mathbf{x}} = \frac{\mathbf{x}}{\|\mathbf{x}\|} \qquad \text{so} \qquad \frac{[3,4]}{5} = [0.6,\, 0.8]

Check: 0.62+0.82=0.36+0.64=10.6^2 + 0.8^2 = 0.36 + 0.64 = 1. The result is a unit vector — pure direction, length exactly one. Normalise both vectors first and their dot product is the cosine similarity, no division needed. That is why vector databases store normalised embeddings.

A matrix is your whole dataset in one object

Stack feature vectors as rows and you have a matrix: a rectangular grid of numbers with mm rows and nn columns, written A∈Rm×nA \in \mathbb{R}^{m \times n}. The entry in row ii, column jj is AijA_{ij}.

In machine learning the convention is fixed and worth burning in: the design matrix XX has nn rows — one per sample — and dd columns — one per feature.

Text
           area   beds   yearhouse 1 [  1400     3    1995 ]house 2 [  1900     4    2004 ]     X is 3 x 3  (n=3 samples, d=3 features)house 3 [  1100     2    1978 ]

Read across a row and you get one house. Read down a column and you get one feature across every house — which is exactly what you need to compute a mean, a standard deviation, or a normalisation.

A few matrices have names because they turn up constantly:

NameShape ruleWhy it matters
Identity IISquare; 1s on the diagonal, 0s elsewhereAI=IA=AAI = IA = A. The "multiply by one" of matrices
DiagonalSquare; non-zero only on the diagonalScales each axis independently; trivial to invert
SymmetricA=ATA = A^T, mirror image across the diagonalCovariance and kernel matrices are always symmetric
ZeroEvery entry 0The additive identity; a common bug when weights fail to initialise

The cheap operations

Addition works entry by entry and needs identical shapes. Scalar multiplication scales every entry. The transpose ATA^T flips rows into columns, so a 2×32 \times 3 becomes a 3×23 \times 2 and AijT=AjiA^T_{ij} = A_{ji}. The Hadamard product A⊙BA \odot B multiplies matching entries and, despite the name, is not matrix multiplication at all — it is what NumPy's * does. Confusing * with @ is the single most common beginner error in numerical Python, and the worst kind, because with square matrices both run happily and give different answers.

Matrix multiplication, properly

This is the operation everything else is built on, so slow down here.

The shape rule. You may multiply AA by BB only when the inner dimensions match:

(m×n)(n×p)=(m×p)(m \times n)(n \times p) = (m \times p)

The nn in the middle must agree and then vanishes. The outer numbers survive as the shape of the result. If the inner numbers disagree, the product simply does not exist — this is not a convention you can bend.

The entry rule. Each output entry is a dot product: row ii of AA with column jj of BB.

Cij=∑k=1nAikBkjC_{ij} = \sum_{k=1}^{n} A_{ik} B_{kj}

Worked in full, with a 2×32 \times 3 times a 3×23 \times 2, which must give a 2×22 \times 2:

A=[123456],B=[789101112]A = \begin{bmatrix} 1 & 2 & 3 \\ 4 & 5 & 6 \end{bmatrix}, \qquad B = \begin{bmatrix} 7 & 8 \\ 9 & 10 \\ 11 & 12 \end{bmatrix}

Take row 1 of AA against column 1 of BB: 1(7)+2(9)+3(11)=7+18+33=581(7) + 2(9) + 3(11) = 7 + 18 + 33 = 58.

Row 1 against column 2: 1(8)+2(10)+3(12)=8+20+36=641(8) + 2(10) + 3(12) = 8 + 20 + 36 = 64.

Row 2 against column 1: 4(7)+5(9)+6(11)=28+45+66=1394(7) + 5(9) + 6(11) = 28 + 45 + 66 = 139.

Row 2 against column 2: 4(8)+5(10)+6(12)=32+50+72=1544(8) + 5(10) + 6(12) = 32 + 50 + 72 = 154.

AB=[5864139154]AB = \begin{bmatrix} 58 & 64 \\ 139 & 154 \end{bmatrix}

What holds, and what famously does not

Matrix multiplication is associative, (AB)C=A(BC)(AB)C = A(BC), and distributive over addition, A(B+C)=AB+ACA(B+C) = AB + AC. It is not commutative. ABAB and BABA are usually different matrices, and often one of them does not even exist. A concrete counterexample:

A=[1234],B=[0110]  ⇒  AB=[2143],BA=[3412]A = \begin{bmatrix} 1 & 2 \\ 3 & 4 \end{bmatrix}, \quad B = \begin{bmatrix} 0 & 1 \\ 1 & 0 \end{bmatrix} \;\Rightarrow\; AB = \begin{bmatrix} 2 & 1 \\ 4 & 3 \end{bmatrix}, \quad BA = \begin{bmatrix} 3 & 4 \\ 1 & 2 \end{bmatrix}

ABAB swapped the columns; BABA swapped the rows. Order encodes meaning, because each matrix is an instruction and instructions do not commute — rotate then stretch is a different thing from stretch then rotate.

One more rule that saves hours of debugging: transposing a product reverses it.

(AB)T=BTAT(AB)^T = B^T A^T

The reversal is forced by shapes. If AA is m×nm \times n and BB is n×pn \times p, then ATBTA^T B^T would be n×mn \times m times p×np \times n, which is illegal unless the numbers coincide by luck. BTATB^T A^T is p×np \times n times n×mn \times m, which works and gives the p×mp \times m shape you want.

If you can state the shape of every array in a line of code before you run it, you will catch most numerical bugs before they happen.

Why batching a model is a matrix multiply

Back to the houses. A linear model has one weight per feature plus an intercept, so for a single house y^=x⋅w+b\hat{y} = \mathbf{x} \cdot \mathbf{w} + b — a dot product and an addition. Now stack all the houses into XX and the whole batch becomes:

y^=Xw+b\hat{\mathbf{y}} = X\mathbf{w} + b

Shapes: XX is n×dn \times d, w\mathbf{w} is d×1d \times 1, so XwX\mathbf{w} is n×1n \times 1 — one prediction per sample, exactly as required. Each row of the output is that row's dot product with the weights, which is precisely the per-house formula. The matrix product is not an approximation of the loop; it is the loop, expressed as one operation the hardware knows how to parallelise.

Numbers, with weights w=[0.15,20]\mathbf{w} = [0.15, 20] on area and bedrooms, and b=50b = 50, prices in thousands:

[140031900411002][0.1520]+50=[210+60285+80165+40]+50=[320415255]\begin{bmatrix} 1400 & 3 \\ 1900 & 4 \\ 1100 & 2 \end{bmatrix} \begin{bmatrix} 0.15 \\ 20 \end{bmatrix} + 50 = \begin{bmatrix} 210 + 60 \\ 285 + 80 \\ 165 + 40 \end{bmatrix} + 50 = \begin{bmatrix} 320 \\ 415 \\ 255 \end{bmatrix}

Every neural network layer is this same shape with a nonlinearity wrapped round it. Change w\mathbf{w} from a vector to a matrix WW of shape d×hd \times h and you get hh outputs per sample instead of one. That is a dense layer.

Doing it in NumPy without the classic mistakes

Python
import numpy as npX = np.array([[1400, 3],              [1900, 4],              [1100, 2]])        # (3, 2)  n=3 samples, d=2 featuresw = np.array([0.15, 20.0])       # (2,)b = 50.0y_hat = X @ w + b                # (3,)  ->  [320., 415., 255.]# @ and np.dot agree for 2-D inputs; @ is the clearer choiceassert np.allclose(X @ w, np.dot(X, w))# Broadcasting: b is a scalar, stretched across all 3 rows for free.# It also works per-column, which is how a bias vector is added in a layer:B = np.array([10.0, -5.0])       # (2,) broadcast down every row of XX_shifted = X + B                # (3, 2), no loop, no copy of B

Broadcasting means NumPy stretches a smaller array across a larger one when the trailing dimensions are compatible, without ever materialising the copies. It saves memory and it silently hides mistakes, so state the shape you expect and assert it.

The error everyone meets at least once:

Python
A = np.random.randn(3, 2)B = np.random.randn(3, 2)A @ B# ValueError: matmul: Input operand 1 has a mismatch in its core dimension 0,# with gufunc signature (n?,k),(k,m?)->(n?,m?) (size 3 is different from 2)

Translated: the inner dimensions are 2 and 3, and they must match. Either you meant A @ B.T (giving 3×33 \times 3), or A.T @ B (giving 2×22 \times 2), or you meant elementwise A * B and reached for the wrong operator. Ask what the output means — a similarity between every pair of samples, or a per-feature accumulation? — and the right form follows.

Almost every shape error is a modelling question in disguise: you have not yet decided what the output is supposed to represent.

Thinking in shapes when you build something

The practical payoff is a habit, not a formula. Before writing a line of numerical code, write down the shape of every array on paper and check the chain multiplies: (n×d)(d×h)→(n×h)(n \times d)(d \times h) \rightarrow (n \times h), then (n×h)(h×1)→(n×1)(n \times h)(h \times 1) \rightarrow (n \times 1). If the chain is legal, the code is probably right. If it is not, no amount of debugging output will help until you fix it.

Three habits that follow from everything above. First, keep samples in rows and features in columns, always — most libraries assume it, and a transposed dataset produces a d×dd \times d matrix that trains without complaint and predicts nonsense. Second, distrust * in any line that should be a matrix product; write @ deliberately and let the shape rule catch you. Third, whenever you are asking "how similar are these two things?", normalise first and then take a dot product, so the answer is about direction rather than about which vector happened to be larger.

Do that consistently and the slow loop never gets written in the first place.