Course Content
LLMs Deep Dive
10 sections · 40 lessons
What is the role of the Jacobian matrix in backpropagation?
What you need to know
A 2 × 2 example
Take f(x1, x2) = (x1 × x2, x1 + x2). Then:
J = [[x2, x1], [ 1, 1]]At x = (2, 3): J = [[3, 2], [1, 1]]Vector–Jacobian products (VJP)
The loss is a single number. Suppose the gradient of the loss with respect to y is v = [1, 0.5]. The gradient with respect to x is:
v^T J = [1x3 + 0.5x1, 1x2 + 0.5x1] = [3.5, 2.5]That is all backpropagation needs from each layer: take v in, give v^T J out. For a matrix multiply y = Wx, the VJP is just W^T v — no Jacobian is ever formed.
Why not build the Jacobian?
One transformer layer maps an 8,192-token sequence of 4,096-wide vectors to another of the same shape: about 33.5 million numbers in and out. Its full Jacobian would have (33.5 million)² ≈ 1.1 × 10¹⁵ entries — far beyond any memory. The VJP costs about the same as the forward pass.
Reverse mode versus forward mode
- Reverse mode (VJP) — one backward pass gives the gradient of one scalar with respect to all inputs. Perfect for training: one loss, billions of parameters.
- Forward mode (JVP) — one pass gives how all outputs change along one input direction. Useful for a few inputs and many outputs.
The softmax Jacobian
For softmax output p, J = diag(p) - p p^T. When p is nearly one-hot, every entry is near zero — the formal reason saturated softmax stops learning.
Where gradients vanish or explode
Backprop multiplies one Jacobian per layer. If each shrinks the gradient by a factor of 0.9, after 50 layers it is 0.9^50 ≈ 0.005. If each stretches by 1.1, it is 1.1^50 ≈ 117. The singular values of the layer Jacobians decide which happens.
A real-life example
A team fine-tuning a model for legal-document summarisation sees the loss suddenly spike to NaN after 3,000 steps. Logging the gradient norm per layer shows it growing from about 1 to over 1,000 in a few steps in the top layers — a product of Jacobians stretching the signal.
They add gradient clipping at a norm of 1.0 (scale the whole gradient down if its length exceeds 1), lower the learning rate, and skip a batch of scanned contracts full of garbled OCR text that triggered huge activations. Training is stable again. Knowing that backprop is a chain of Jacobian products told them to look at per-layer norms rather than just the loss curve.
Follow-up questions to expect
- "How does PyTorch compute gradients?" — It records the operations in the forward pass, then walks them in reverse, calling each operation's VJP rule.
- "When would you compute a full Jacobian?" — For small functions: sensitivity analysis, some physics models, or research on training dynamics, using
torch.func.jacrev. - "What is the Hessian?" — The matrix of second derivatives, the Jacobian of the gradient. It describes curvature; like the Jacobian, it is used through products, not built.