Course Content
Mathematics for Machine Learning
5 sections · 13 lessons
Vectors, Matrices, and Operations
You have a price model for houses. It is simple: each listing has 50 numeric features — floor area, bedroom count, year built and so on — and you multiply each one by its own weight, add a baseline, and out comes a predicted price. You have a million listings to score. So you write the obvious thing.
1predictions = []2for listing in listings: # 1,000,000 listings, 50 features each3 total = intercept4 for feature, weight in zip(listing, weights):5 total += feature * weight6 predictions.append(total)It works. It also takes five or six seconds, even on a fast laptop. That is fine once, and painful when you are tuning the model and want to re-score every time you change a weight. Now watch the same computation written differently:
predictions = X @ w + b # done in about 7 millisecondsSame arithmetic. Same answers. Several hundred times faster, because that one line hands the whole job to a compiled routine that keeps the CPU's vector units busy and never touches the Python interpreter. That gap is not a trick. It is the reason machine learning is written in the language of vectors and matrices rather than loops. Learn to see your data as arrays of numbers and the operations on them as a handful of named products, and you get speed, clarity, and a vocabulary shared by every ML library ever written. Here is that vocabulary, from the bottom.
A vector is a list of numbers that means something
A vector is an ordered list of numbers. That is the whole definition. What makes it useful is that the order carries meaning: slot one is always floor area, slot two is always bedrooms, and every vector in your dataset agrees on that.
This is one house: 1400 square feet, 3 bedrooms, built in 1995. Because it has three numbers we say it lives in R3 — three-dimensional real space. A vector with 768 numbers lives in R768, which you cannot picture and never need to.
There are three ways to read the same object, and switching between them freely is most of the skill:
| Reading | What it says | When it helps |
|---|---|---|
| A list | Three stored numbers, nothing more | Writing code, thinking about memory and shapes |
| A point | A location in space, one axis per feature | Distance, clustering, nearest neighbours |
| An arrow | A direction plus a magnitude, from the origin | Gradients, similarity, anything about angle |
In machine learning the word you will meet is feature vector: one row of your dataset, one example, with every measurement about it laid out in a fixed order. A house, a customer, a sentence turned into 768 numbers by a language model — all the same shape of object.
Adding vectors and scaling them
Two vectors of the same length add position by position:
Multiplying by a single number — a scalar — stretches every component by the same factor: 2⋅[3,4]=[6,8]. The arrow keeps its direction and doubles in length. A negative scalar flips it round. These two operations are the entire content of gradient descent: your weights are a vector, the gradient is a vector, and the update w←w−η∇ is a scalar multiply followed by a subtraction.
The dot product, and why it is really about angle
The dot product multiplies matching components and sums the results:
With a=[3,4] and b=[1,2]: 3×1+4×2=3+8=11. One number out, no matter how long the vectors are.
That arithmetic hides something geometric. The same quantity equals
where ∥⋅∥ means length and θ is the angle between the two arrows. In plain English: the dot product is big when the vectors are long and point the same way, zero when they are at right angles, and negative when they point apart. Check it on our numbers. ∥a∥=9+16=5 and ∥b∥=1+4≈2.236, so cosθ=11/(5×2.236)=0.984, giving θ≈10.3∘. Nearly the same direction, which matches the picture.
Rearrange that formula and you have cosine similarity, the standard way to ask "are these two things alike?":
Dividing by both lengths strips out magnitude and leaves only direction, on a scale from −1 to 1. This is why search engines and recommendation systems compare embeddings with cosine similarity rather than raw dot products: a long document should not beat a short one simply for being long.
The dot product is the only bridge between arithmetic and geometry you need: multiply-and-sum on one side, angle on the other.
Length, and why models care about it
The length of a vector — its norm, written ∥x∥ — is Pythagoras in as many dimensions as you like:
For [3,4] that is 9+16=5. This is the L2 norm, the everyday notion of distance. Its sibling the L1 norm adds absolute values instead, ∥[3,4]∥1=3+4=7 — the distance you would walk on a street grid. The choice matters in regularisation: penalising the L2 norm of the weights shrinks them all smoothly, while penalising the L1 norm drives many of them to exactly zero and gives you feature selection for free.
To normalise a vector, divide it by its own length:
Check: 0.62+0.82=0.36+0.64=1. The result is a unit vector — pure direction, length exactly one. Normalise both vectors first and their dot product is the cosine similarity, no division needed. That is why vector databases store normalised embeddings.
A matrix is your whole dataset in one object
Stack feature vectors as rows and you have a matrix: a rectangular grid of numbers with m rows and n columns, written A∈Rm×n. The entry in row i, column j is Aij.
In machine learning the convention is fixed and worth burning in: the design matrix X has n rows — one per sample — and d columns — one per feature.
area beds yearhouse 1 [ 1400 3 1995 ]house 2 [ 1900 4 2004 ] X is 3 x 3 (n=3 samples, d=3 features)house 3 [ 1100 2 1978 ]Read across a row and you get one house. Read down a column and you get one feature across every house — which is exactly what you need to compute a mean, a standard deviation, or a normalisation.
A few matrices have names because they turn up constantly:
| Name | Shape rule | Why it matters |
|---|---|---|
| Identity I | Square; 1s on the diagonal, 0s elsewhere | AI=IA=A. The "multiply by one" of matrices |
| Diagonal | Square; non-zero only on the diagonal | Scales each axis independently; trivial to invert |
| Symmetric | A=AT, mirror image across the diagonal | Covariance and kernel matrices are always symmetric |
| Zero | Every entry 0 | The additive identity; a common bug when weights fail to initialise |
The cheap operations
Addition works entry by entry and needs identical shapes. Scalar multiplication scales every entry. The transpose AT flips rows into columns, so a 2×3 becomes a 3×2 and AijT=Aji. The Hadamard product A⊙B multiplies matching entries and, despite the name, is not matrix multiplication at all — it is what NumPy's * does. Confusing * with @ is the single most common beginner error in numerical Python, and the worst kind, because with square matrices both run happily and give different answers.
Matrix multiplication, properly
This is the operation everything else is built on, so slow down here.
The shape rule. You may multiply A by B only when the inner dimensions match:
The n in the middle must agree and then vanishes. The outer numbers survive as the shape of the result. If the inner numbers disagree, the product simply does not exist — this is not a convention you can bend.
The entry rule. Each output entry is a dot product: row i of A with column j of B.
Worked in full, with a 2×3 times a 3×2, which must give a 2×2:
Take row 1 of A against column 1 of B: 1(7)+2(9)+3(11)=7+18+33=58.
Row 1 against column 2: 1(8)+2(10)+3(12)=8+20+36=64.
Row 2 against column 1: 4(7)+5(9)+6(11)=28+45+66=139.
Row 2 against column 2: 4(8)+5(10)+6(12)=32+50+72=154.
What holds, and what famously does not
Matrix multiplication is associative, (AB)C=A(BC), and distributive over addition, A(B+C)=AB+AC. It is not commutative. AB and BA are usually different matrices, and often one of them does not even exist. A concrete counterexample:
AB swapped the columns; BA swapped the rows. Order encodes meaning, because each matrix is an instruction and instructions do not commute — rotate then stretch is a different thing from stretch then rotate.
One more rule that saves hours of debugging: transposing a product reverses it.
The reversal is forced by shapes. If A is m×n and B is n×p, then ATBT would be n×m times p×n, which is illegal unless the numbers coincide by luck. BTAT is p×n times n×m, which works and gives the p×m shape you want.
If you can state the shape of every array in a line of code before you run it, you will catch most numerical bugs before they happen.
Why batching a model is a matrix multiply
Back to the houses. A linear model has one weight per feature plus an intercept, so for a single house y^=x⋅w+b — a dot product and an addition. Now stack all the houses into X and the whole batch becomes:
Shapes: X is n×d, w is d×1, so Xw is n×1 — one prediction per sample, exactly as required. Each row of the output is that row's dot product with the weights, which is precisely the per-house formula. The matrix product is not an approximation of the loop; it is the loop, expressed as one operation the hardware knows how to parallelise.
Numbers, with weights w=[0.15,20] on area and bedrooms, and b=50, prices in thousands:
Every neural network layer is this same shape with a nonlinearity wrapped round it. Change w from a vector to a matrix W of shape d×h and you get h outputs per sample instead of one. That is a dense layer.
Doing it in NumPy without the classic mistakes
1import numpy as np23X = np.array([[1400, 3],4 [1900, 4],5 [1100, 2]]) # (3, 2) n=3 samples, d=2 features6w = np.array([0.15, 20.0]) # (2,)7b = 50.089y_hat = X @ w + b # (3,) -> [320., 415., 255.]1011# @ and np.dot agree for 2-D inputs; @ is the clearer choice12assert np.allclose(X @ w, np.dot(X, w))1314# Broadcasting: b is a scalar, stretched across all 3 rows for free.15# It also works per-column, which is how a bias vector is added in a layer:16B = np.array([10.0, -5.0]) # (2,) broadcast down every row of X17X_shifted = X + B # (3, 2), no loop, no copy of BBroadcasting means NumPy stretches a smaller array across a larger one when the trailing dimensions are compatible, without ever materialising the copies. It saves memory and it silently hides mistakes, so state the shape you expect and assert it.
The error everyone meets at least once:
1A = np.random.randn(3, 2)2B = np.random.randn(3, 2)3A @ B4# ValueError: matmul: Input operand 1 has a mismatch in its core dimension 0,5# with gufunc signature (n?,k),(k,m?)->(n?,m?) (size 3 is different from 2)Translated: the inner dimensions are 2 and 3, and they must match. Either you meant A @ B.T (giving 3×3), or A.T @ B (giving 2×2), or you meant elementwise A * B and reached for the wrong operator. Ask what the output means — a similarity between every pair of samples, or a per-feature accumulation? — and the right form follows.
Almost every shape error is a modelling question in disguise: you have not yet decided what the output is supposed to represent.
Thinking in shapes when you build something
The practical payoff is a habit, not a formula. Before writing a line of numerical code, write down the shape of every array on paper and check the chain multiplies: (n×d)(d×h)→(n×h), then (n×h)(h×1)→(n×1). If the chain is legal, the code is probably right. If it is not, no amount of debugging output will help until you fix it.
Three habits that follow from everything above. First, keep samples in rows and features in columns, always — most libraries assume it, and a transposed dataset produces a d×d matrix that trains without complaint and predicts nonsense. Second, distrust * in any line that should be a matrix product; write @ deliberately and let the shape rule catch you. Third, whenever you are asking "how similar are these two things?", normalise first and then take a dot product, so the answer is about direction rather than about which vector happened to be larger.
Do that consistently and the slow loop never gets written in the first place.