Course Content
Deep Learning with TensorFlow and PyTorch
4 sections · 15 lessons
Visualizing Training with TensorBoard
Your training script prints one line per epoch. After 200 epochs you have 200 lines. You change the learning rate and run it again: 400 lines. You try six architectures overnight: 1,400 lines of numbers in a terminal buffer, and the question you actually want to answer is "which of these six started overfitting earliest, and at what epoch".
You cannot answer it from that scrollback. You could paste the numbers into a spreadsheet, which people do, and which takes twenty minutes per experiment and is abandoned by the third day.
Worse, there are questions the print statement cannot answer at any cost. Are the gradients in layer 3 shrinking over training? What do the images that the model gets wrong have in common? Have the weights in the final layer stopped moving? Each of those requires seeing a distribution or an image, not a scalar.
TensorBoard is a small web application that reads structured logs your training script writes and renders them as interactive charts, histograms, images and projections. It costs one line for the basic case and it changes what you are able to notice.
Getting it running
1pip install tensorboard23# Point it at the directory your script writes to, then open the URL it prints4tensorboard --logdir=runs5# TensorBoard 2.x at http://localhost:6006/67# Inside a Jupyter or Colab notebook instead:8# %load_ext tensorboard9# %tensorboard --logdir runsIt works with both frameworks. PyTorch writes the same event-file format through torch.utils.tensorboard, so nothing about the tool is TensorFlow-specific any more.
TensorFlow and Keras: one callback
1import tensorflow as tf2from datetime import datetime34log_dir = "runs/" + datetime.now().strftime("%Y%m%d-%H%M%S")5tb = tf.keras.callbacks.TensorBoard(6 log_dir=log_dir,7 histogram_freq=1, # weight/bias distributions every epoch8 write_graph=True, # the model architecture diagram9 update_freq="epoch", # "batch" for finer resolution, larger logs10 profile_batch=(10, 20), # profile these batches for a performance trace11)1213model.fit(train_ds, validation_data=val_ds, epochs=50, callbacks=[tb])The timestamped directory matters. Every run gets its own subdirectory, TensorBoard treats each subdirectory as a separate series, and you can overlay them on the same axes. Reuse the same directory and two runs interleave into one nonsensical curve.
When you need something fit() does not log, write it yourself:
1writer = tf.summary.create_file_writer(log_dir)23with writer.as_default():4 for epoch in range(epochs):5 loss = train_one_epoch()6 tf.summary.scalar("loss/train", loss, step=epoch)7 tf.summary.scalar("learning_rate",8 optimizer.learning_rate.numpy(), step=epoch)9 for var in model.trainable_variables:10 tf.summary.histogram(f"weights/{var.path}", var, step=epoch) # .path is unique; .name is just "kernel"11 writer.flush()PyTorch: SummaryWriter
1from torch.utils.tensorboard import SummaryWriter2from datetime import datetime34writer = SummaryWriter(f"runs/{datetime.now():%Y%m%d-%H%M%S}_lr1e-3_bs64")56global_step = 07for epoch in range(epochs):8 model.train()9 for xb, yb in train_loader:10 xb, yb = xb.to(device), yb.to(device)11 opt.zero_grad()12 loss = loss_fn(model(xb), yb)13 loss.backward()14 opt.step()1516 if global_step % 50 == 0: # not every step17 writer.add_scalar("loss/train", loss.item(), global_step)18 writer.add_scalar("lr", opt.param_groups[0]["lr"], global_step)19 global_step += 12021 val_loss, val_acc = evaluate(model, val_loader)22 writer.add_scalar("loss/val", val_loss, epoch)23 writer.add_scalar("accuracy/val", val_acc, epoch)2425 # Distributions of weights and their gradients, once per epoch26 for name, p in model.named_parameters():27 writer.add_histogram(f"weights/{name}", p, epoch)28 if p.grad is not None:29 writer.add_histogram(f"grads/{name}", p.grad, epoch)3031writer.close() # flushes buffered events -- charts look truncated without itTwo conventions in that code are worth adopting. The tag names use a slash — loss/train, loss/val — and TensorBoard groups tags by the part before the slash, so both loss curves land on one chart where you can see the gap between them. And the run directory name encodes the hyperparameters, so the legend tells you what each line is without cross-referencing a notebook.
You can also log the architecture itself:
example_batch, _ = next(iter(train_loader))writer.add_graph(model, example_batch.to(device))What to log, and how often
| Quantity | Frequency | What it tells you |
|---|---|---|
| Training loss | Every 20–100 steps | Whether optimisation is progressing; instability |
| Validation loss and metric | Every epoch | Generalisation; when to stop |
| Learning rate | Every epoch | Confirms your schedule is doing what you think |
| Gradient norms | Every epoch | Vanishing or exploding gradients |
| Weight histograms | Every epoch (or every 5) | Whether layers are still learning; dead units |
| Sample predictions or errors | Every 5–10 epochs | What kind of mistakes the model makes |
| Embeddings | Once, at the end | Whether the learned representation separates classes |
Logging every step is a real cost. Histograms are the expensive item — they serialise every parameter — and logging them per step on a large model can double your training time and fill a disk with gigabytes of event files. Scalars are cheap; histograms and images are not.
Reading histograms, which is where the value hides
The histogram view shows the distribution of a tensor's values, with one ridge per epoch stacked into the depth of the plot. Scalars tell you that something is wrong; histograms usually tell you what.
| What the shape does over epochs | Meaning |
|---|---|
| Weights start narrow, gradually widen | Normal, healthy learning |
| Weights identical from epoch 1 to epoch 50 | That layer is frozen or receives no gradient |
| Weights spread wider and wider without limit | No regularisation; consider weight decay |
| Gradient histogram collapses to a spike at zero | Vanishing gradients in that layer |
| Gradient histogram has a very long tail | Exploding gradients; add clipping |
| Activation histogram is a spike at exactly 0 | Dead ReLU units |
| Activation histogram piles up at ±1 | Saturated tanh or sigmoid; gradients are dying |
The last two require logging activations, which needs a forward hook:
1acts = {}2def capture(name):3 return lambda module, inp, out: acts.__setitem__(name, out.detach())45for name, layer in model.named_modules():6 if isinstance(layer, torch.nn.ReLU):7 layer.register_forward_hook(capture(name))89model.eval()10with torch.no_grad():11 model(example_batch.to(device))12for name, a in acts.items():13 writer.add_histogram(f"activations/{name}", a, epoch)14 writer.add_scalar(f"dead_fraction/{name}",15 (a == 0).all(dim=0).float().mean().item(), epoch)That dead_fraction scalar is one of the highest-value things you can log for a ReLU network. It is a single number per layer, cheap to compute, and it turns "the model underperforms and I don't know why" into "layer 4 has 60% dead units, so my learning rate is too high".
Images: seeing what the model sees
1import torchvision23# A grid of the current batch, after augmentation -- verifies your pipeline4grid = torchvision.utils.make_grid(xb[:16], normalize=True, nrow=4)5writer.add_image("inputs/augmented", grid, epoch)Logging augmented inputs catches a whole class of silent bugs: normalisation applied twice, a channel order swapped, a crop that removes the object of interest. A picture makes those obvious in a second; no scalar ever will.
More valuable still is logging the model's mistakes:
1model.eval()2with torch.no_grad():3 logits = model(xb.to(device))4 preds = logits.argmax(1).cpu()5wrong = (preds != yb).nonzero(as_tuple=True)[0][:16]6if len(wrong):7 writer.add_image("errors/worst",8 torchvision.utils.make_grid(xb[wrong], normalize=True),9 epoch)10 writer.add_text("errors/labels",11 ", ".join(f"true={yb[i]} pred={preds[i]}" for i in wrong),12 epoch)Twenty misclassified images usually reveal a pattern in under a minute — every failure is a dark photograph, or every failure is a particular breed, or every failure has the object at the edge of the frame. That pattern points at a fix. An aggregate accuracy number points nowhere.
A confusion matrix as an image
1import io, matplotlib.pyplot as plt, numpy as np2from sklearn.metrics import confusion_matrix3from PIL import Image4import torchvision.transforms.functional as TF56def log_confusion(writer, y_true, y_pred, class_names, epoch):7 cm = confusion_matrix(y_true, y_pred, normalize="true")8 fig, ax = plt.subplots(figsize=(6, 6))9 ax.imshow(cm, cmap="Blues", vmin=0, vmax=1)10 ax.set_xticks(range(len(class_names)), class_names, rotation=90)11 ax.set_yticks(range(len(class_names)), class_names)12 ax.set_xlabel("predicted"); ax.set_ylabel("true")13 for i in range(len(class_names)):14 for j in range(len(class_names)):15 ax.text(j, i, f"{cm[i, j]:.2f}", ha="center", fontsize=7)1617 buf = io.BytesIO()18 fig.savefig(buf, format="png", bbox_inches="tight", dpi=110)19 plt.close(fig)20 buf.seek(0)21 writer.add_image("confusion_matrix",22 TF.to_tensor(Image.open(buf).convert("RGB")), epoch)Watching the confusion matrix evolve across epochs shows which class pairs the model resolves first and which it never separates. Two classes that stay confused with each other after fifty epochs are usually genuinely ambiguous in the data, and the fix is in the labels, not the architecture.
The embedding projector
Take the activations of the penultimate layer — the representation the classifier head sees — and project them to three dimensions.
1features, labels, images = [], [], []2model.eval()3with torch.no_grad():4 for xb, yb in val_loader:5 feats = model.backbone(xb.to(device)) # before the final layer6 features.append(feats.cpu()); labels.append(yb); images.append(xb)7 if len(labels) * xb.size(0) >= 1000: # 1000 points is plenty8 break910writer.add_embedding(11 torch.cat(features), metadata=torch.cat(labels).tolist(),12 label_img=torch.cat(images), tag="penultimate")Well-separated clusters, one per class, mean the network has learned a representation where classification is easy. Overlapping clouds mean it has not, and no amount of tuning the final layer will rescue it. Points of one class sitting inside another class's cluster are worth hovering over — they are frequently mislabelled examples, and finding a handful of those can be worth more than a week of architecture search.
Comparing runs
The whole point of separate run directories is overlaying them:
runs/ 20260518-0912_lr1e-2_bs32/ 20260518-0940_lr1e-3_bs32/ 20260518-1015_lr1e-4_bs32/ 20260518-1102_lr1e-3_bs128/Start TensorBoard on runs/ and all four appear on the same axes, with the directory name as the legend. For a systematic sweep, the HParams dashboard turns this into a sortable table:
1from torch.utils.tensorboard import SummaryWriter23for lr in [1e-2, 1e-3, 1e-4]:4 for dropout in [0.0, 0.3, 0.5]:5 w = SummaryWriter(f"runs/hp_lr{lr}_do{dropout}")6 best_acc = train_model(lr=lr, dropout=dropout, writer=w)7 w.add_hparams({"lr": lr, "dropout": dropout},8 {"hparam/best_val_acc": best_acc})9 w.close()The dashboard then gives you a parallel-coordinates view showing which hyperparameters actually correlate with the metric. Frequently the answer is "only the learning rate", which saves you from tuning the other five.
Diagnosing from the curves
| Shape | Diagnosis | Action |
|---|---|---|
| Train and val both falling together | Healthy | Keep going |
| Train falls, val flattens then rises | Overfitting from the turning point | Early stopping, more regularisation, more data |
| Both flat and high from the start | Underfitting, or the learning rate is far too low | Bigger model; raise the LR by 10× |
| Loss spikes then recovers repeatedly | Learning rate slightly too high | Add a decay schedule |
Loss jumps to nan | Exploding gradients or a numerical bug | Clip gradients; check for log of zero |
| Sudden drop exactly when LR is cut | Normal — the optimiser can finally settle | Nothing; confirms the schedule works |
| Val loss noticeably below train loss | Usually just dropout inflating the train figure | Nothing, if the gap is small and stable |
Making it a habit rather than a chore
Instrument the run before you start it, not after it disappoints you. Re-running a six-hour job because you did not log gradient norms is a bad trade against the ninety seconds it takes to add the lines.
Encode hyperparameters in the run directory name. Six months from now, runs/exp7 tells you nothing and runs/20260518-1015_resnet18_lr1e-3_bs128_aug tells you everything. If you also write the full config as a JSON file into the same directory, the run becomes genuinely reproducible.
Keep a small fixed set of tags across all your projects — loss/train, loss/val, lr, grad_norm/total, dead_fraction/* — so that comparing an experiment from today with one from last month does not require translating between two logging schemes.
Finally, be clear about what TensorBoard is for. It is an inspection tool for the run you are looking at, not an experiment database. It has no notion of code version, dataset version or environment, and its logs get deleted when someone cleans up the disk. Once you are running dozens of experiments that matter, pair it with something that records the git commit, the config and the data hash alongside the metrics. TensorBoard answers "what is this model doing right now"; that is a different question from "what exactly produced the model we shipped in March", and you will eventually need both.