LLMs Deep Dive

Course Content

LLMs Deep Dive

10 sections · 40 lessons

What is the role of the Jacobian matrix in backpropagation?


What you need to know

A 2 × 2 example

Take f(x1, x2) = (x1 × x2, x1 + x2). Then:

Text
J = [[x2, x1],     [ 1,  1]]At x = (2, 3):  J = [[3, 2],                     [1, 1]]

Vector–Jacobian products (VJP)

The loss is a single number. Suppose the gradient of the loss with respect to y is v = [1, 0.5]. The gradient with respect to x is:

Text
v^T J = [1x3 + 0.5x1,  1x2 + 0.5x1] = [3.5, 2.5]

That is all backpropagation needs from each layer: take v in, give v^T J out. For a matrix multiply y = Wx, the VJP is just W^T v — no Jacobian is ever formed.

Why not build the Jacobian?

One transformer layer maps an 8,192-token sequence of 4,096-wide vectors to another of the same shape: about 33.5 million numbers in and out. Its full Jacobian would have (33.5 million)² ≈ 1.1 × 10¹⁵ entries — far beyond any memory. The VJP costs about the same as the forward pass.

Reverse mode versus forward mode

  • Reverse mode (VJP) — one backward pass gives the gradient of one scalar with respect to all inputs. Perfect for training: one loss, billions of parameters.
  • Forward mode (JVP) — one pass gives how all outputs change along one input direction. Useful for a few inputs and many outputs.

The softmax Jacobian

For softmax output p, J = diag(p) - p p^T. When p is nearly one-hot, every entry is near zero — the formal reason saturated softmax stops learning.

Where gradients vanish or explode

Backprop multiplies one Jacobian per layer. If each shrinks the gradient by a factor of 0.9, after 50 layers it is 0.9^50 ≈ 0.005. If each stretches by 1.1, it is 1.1^50 ≈ 117. The singular values of the layer Jacobians decide which happens.

A real-life example

A team fine-tuning a model for legal-document summarisation sees the loss suddenly spike to NaN after 3,000 steps. Logging the gradient norm per layer shows it growing from about 1 to over 1,000 in a few steps in the top layers — a product of Jacobians stretching the signal.

They add gradient clipping at a norm of 1.0 (scale the whole gradient down if its length exceeds 1), lower the learning rate, and skip a batch of scanned contracts full of garbled OCR text that triggered huge activations. Training is stable again. Knowing that backprop is a chain of Jacobian products told them to look at per-layer norms rather than just the loss curve.

Follow-up questions to expect

  • "How does PyTorch compute gradients?" — It records the operations in the forward pass, then walks them in reverse, calling each operation's VJP rule.
  • "When would you compute a full Jacobian?" — For small functions: sensitivity analysis, some physics models, or research on training dynamics, using torch.func.jacrev.
  • "What is the Hessian?" — The matrix of second derivatives, the Jacobian of the gradient. It describes curvature; like the Jacobian, it is used through products, not built.