LLMs Deep Dive

Course Content

LLMs Deep Dive

10 sections · 40 lessons

What is cross-entropy loss and why is it used in language models?


What you need to know

Text
Loss at one position = -log p(correct token)

Worked numbers

p(correct token)Loss (natural log)
0.90.105
0.50.693
0.12.303
0.014.605

Going from 0.9 to 0.5 costs about 0.6. Going from 0.1 to 0.01 costs another 2.3. As p approaches 0, loss goes to infinity. The model is pushed hardest where it is confidently wrong — exactly where it needs to learn.

The gradient is simple

With softmax outputs p and a one-hot target y, the gradient of the loss with respect to the logits is:

Text
gradient = p - y

Example: probabilities [0.7, 0.2, 0.1], correct token is index 1. Gradient = [0.7, -0.8, 0.1]. Push the correct logit up by a lot (0.8), push the wrong ones down in proportion to how much probability they stole. No vanishing, no complicated terms. Mean squared error on softmax outputs gives a weaker, messier gradient, which is why it is not used for classification.

In code

Python
import torchimport torch.nn.functional as Flogits = torch.tensor([[2.0, 1.0, 0.1]])   # one position, 3-token vocabtarget = torch.tensor([0])                  # correct token is index 0print(F.cross_entropy(logits, target))      # tensor(0.4170) = -log(0.659)

F.cross_entropy applies log-softmax and the negative log-likelihood in one stable step, so you pass raw logits, not probabilities.

Two practical details

  • Loss masking in fine-tuning — for chat data you usually compute loss only on the assistant's tokens, not the user's prompt, so the model learns to answer rather than to imitate customers.
  • Label smoothing — the target puts, say, 0.9 on the correct token and spreads 0.1 over others, which reduces over-confidence. Common in translation, less so in LLM pretraining.

A real-life example

A bank fine-tunes an open model on 30,000 past support chats. Before training, validation loss on assistant tokens is 1.9 (perplexity about 6.7). After two epochs it is 1.4 (perplexity about 4.1): the model is much less "surprised" by how the bank's agents actually reply.

Then the team spots a problem: loss on chats in Hinglish is 2.6, far worse than English at 1.3. Averaging hid it. They add 5,000 more Hinglish chats, and report loss per language from then on. They also remember that lower loss means the model imitates the data better, not that its answers are correct, so they still run a task evaluation on 300 real questions.

Follow-up questions to expect

  • "What is the relationship between cross-entropy and KL divergence?" — Cross-entropy(P, Q) = entropy(P) + KL(P‖Q). With a fixed target P, minimising cross-entropy is the same as minimising KL.
  • "Why log base e and not 2?" — Either works; base 2 gives bits, base e gives nats. Frameworks use natural log.
  • "Does lower perplexity always mean a better chatbot?" — No. It measures next-token prediction on some text; helpfulness and correctness need task evaluations.