Deep Learning with TensorFlow and PyTorch

Visualizing Training with TensorBoard


Your training script prints one line per epoch. After 200 epochs you have 200 lines. You change the learning rate and run it again: 400 lines. You try six architectures overnight: 1,400 lines of numbers in a terminal buffer, and the question you actually want to answer is "which of these six started overfitting earliest, and at what epoch".

You cannot answer it from that scrollback. You could paste the numbers into a spreadsheet, which people do, and which takes twenty minutes per experiment and is abandoned by the third day.

Worse, there are questions the print statement cannot answer at any cost. Are the gradients in layer 3 shrinking over training? What do the images that the model gets wrong have in common? Have the weights in the final layer stopped moving? Each of those requires seeing a distribution or an image, not a scalar.

TensorBoard is a small web application that reads structured logs your training script writes and renders them as interactive charts, histograms, images and projections. It costs one line for the basic case and it changes what you are able to notice.

Five things worth logging, and what each answersOne run, five viewsScalars — isloss still falling?Histograms — areweights drifting?Images — what doesit actually see?Confusion matrix —which class fails?Embeddings — doclasses separate?
Scalars tell you a run went wrong; histograms are usually the panel that tells you why.

Getting it running

Bash
pip install tensorboard# Point it at the directory your script writes to, then open the URL it printstensorboard --logdir=runs# TensorBoard 2.x at http://localhost:6006/# Inside a Jupyter or Colab notebook instead:# %load_ext tensorboard# %tensorboard --logdir runs

It works with both frameworks. PyTorch writes the same event-file format through torch.utils.tensorboard, so nothing about the tool is TensorFlow-specific any more.

TensorFlow and Keras: one callback

Python
import tensorflow as tffrom datetime import datetimelog_dir = "runs/" + datetime.now().strftime("%Y%m%d-%H%M%S")tb = tf.keras.callbacks.TensorBoard(    log_dir=log_dir,    histogram_freq=1,        # weight/bias distributions every epoch    write_graph=True,        # the model architecture diagram    update_freq="epoch",     # "batch" for finer resolution, larger logs    profile_batch=(10, 20),  # profile these batches for a performance trace)model.fit(train_ds, validation_data=val_ds, epochs=50, callbacks=[tb])

The timestamped directory matters. Every run gets its own subdirectory, TensorBoard treats each subdirectory as a separate series, and you can overlay them on the same axes. Reuse the same directory and two runs interleave into one nonsensical curve.

When you need something fit() does not log, write it yourself:

Python
writer = tf.summary.create_file_writer(log_dir)with writer.as_default():    for epoch in range(epochs):        loss = train_one_epoch()        tf.summary.scalar("loss/train", loss, step=epoch)        tf.summary.scalar("learning_rate",                          optimizer.learning_rate.numpy(), step=epoch)        for var in model.trainable_variables:            tf.summary.histogram(f"weights/{var.path}", var, step=epoch)   # .path is unique; .name is just "kernel"        writer.flush()

PyTorch: SummaryWriter

Python
from torch.utils.tensorboard import SummaryWriterfrom datetime import datetimewriter = SummaryWriter(f"runs/{datetime.now():%Y%m%d-%H%M%S}_lr1e-3_bs64")global_step = 0for epoch in range(epochs):    model.train()    for xb, yb in train_loader:        xb, yb = xb.to(device), yb.to(device)        opt.zero_grad()        loss = loss_fn(model(xb), yb)        loss.backward()        opt.step()        if global_step % 50 == 0:                       # not every step            writer.add_scalar("loss/train", loss.item(), global_step)            writer.add_scalar("lr", opt.param_groups[0]["lr"], global_step)        global_step += 1    val_loss, val_acc = evaluate(model, val_loader)    writer.add_scalar("loss/val", val_loss, epoch)    writer.add_scalar("accuracy/val", val_acc, epoch)    # Distributions of weights and their gradients, once per epoch    for name, p in model.named_parameters():        writer.add_histogram(f"weights/{name}", p, epoch)        if p.grad is not None:            writer.add_histogram(f"grads/{name}", p.grad, epoch)writer.close()      # flushes buffered events -- charts look truncated without it

Two conventions in that code are worth adopting. The tag names use a slash — loss/train, loss/val — and TensorBoard groups tags by the part before the slash, so both loss curves land on one chart where you can see the gap between them. And the run directory name encodes the hyperparameters, so the legend tells you what each line is without cross-referencing a notebook.

You can also log the architecture itself:

Python
example_batch, _ = next(iter(train_loader))writer.add_graph(model, example_batch.to(device))

What to log, and how often

QuantityFrequencyWhat it tells you
Training lossEvery 20–100 stepsWhether optimisation is progressing; instability
Validation loss and metricEvery epochGeneralisation; when to stop
Learning rateEvery epochConfirms your schedule is doing what you think
Gradient normsEvery epochVanishing or exploding gradients
Weight histogramsEvery epoch (or every 5)Whether layers are still learning; dead units
Sample predictions or errorsEvery 5–10 epochsWhat kind of mistakes the model makes
EmbeddingsOnce, at the endWhether the learned representation separates classes

Logging every step is a real cost. Histograms are the expensive item — they serialise every parameter — and logging them per step on a large model can double your training time and fill a disk with gigabytes of event files. Scalars are cheap; histograms and images are not.

Reading histograms, which is where the value hides

The histogram view shows the distribution of a tensor's values, with one ridge per epoch stacked into the depth of the plot. Scalars tell you that something is wrong; histograms usually tell you what.

What the shape does over epochsMeaning
Weights start narrow, gradually widenNormal, healthy learning
Weights identical from epoch 1 to epoch 50That layer is frozen or receives no gradient
Weights spread wider and wider without limitNo regularisation; consider weight decay
Gradient histogram collapses to a spike at zeroVanishing gradients in that layer
Gradient histogram has a very long tailExploding gradients; add clipping
Activation histogram is a spike at exactly 0Dead ReLU units
Activation histogram piles up at ±1Saturated tanh or sigmoid; gradients are dying

The last two require logging activations, which needs a forward hook:

Python
acts = {}def capture(name):    return lambda module, inp, out: acts.__setitem__(name, out.detach())for name, layer in model.named_modules():    if isinstance(layer, torch.nn.ReLU):        layer.register_forward_hook(capture(name))model.eval()with torch.no_grad():    model(example_batch.to(device))for name, a in acts.items():    writer.add_histogram(f"activations/{name}", a, epoch)    writer.add_scalar(f"dead_fraction/{name}",                      (a == 0).all(dim=0).float().mean().item(), epoch)

That dead_fraction scalar is one of the highest-value things you can log for a ReLU network. It is a single number per layer, cheap to compute, and it turns "the model underperforms and I don't know why" into "layer 4 has 60% dead units, so my learning rate is too high".

Images: seeing what the model sees

Python
import torchvision# A grid of the current batch, after augmentation -- verifies your pipelinegrid = torchvision.utils.make_grid(xb[:16], normalize=True, nrow=4)writer.add_image("inputs/augmented", grid, epoch)

Logging augmented inputs catches a whole class of silent bugs: normalisation applied twice, a channel order swapped, a crop that removes the object of interest. A picture makes those obvious in a second; no scalar ever will.

More valuable still is logging the model's mistakes:

Python
model.eval()with torch.no_grad():    logits = model(xb.to(device))    preds = logits.argmax(1).cpu()wrong = (preds != yb).nonzero(as_tuple=True)[0][:16]if len(wrong):    writer.add_image("errors/worst",                     torchvision.utils.make_grid(xb[wrong], normalize=True),                     epoch)    writer.add_text("errors/labels",                    ", ".join(f"true={yb[i]} pred={preds[i]}" for i in wrong),                    epoch)

Twenty misclassified images usually reveal a pattern in under a minute — every failure is a dark photograph, or every failure is a particular breed, or every failure has the object at the edge of the frame. That pattern points at a fix. An aggregate accuracy number points nowhere.

A confusion matrix as an image

Python
import io, matplotlib.pyplot as plt, numpy as npfrom sklearn.metrics import confusion_matrixfrom PIL import Imageimport torchvision.transforms.functional as TFdef log_confusion(writer, y_true, y_pred, class_names, epoch):    cm = confusion_matrix(y_true, y_pred, normalize="true")    fig, ax = plt.subplots(figsize=(6, 6))    ax.imshow(cm, cmap="Blues", vmin=0, vmax=1)    ax.set_xticks(range(len(class_names)), class_names, rotation=90)    ax.set_yticks(range(len(class_names)), class_names)    ax.set_xlabel("predicted"); ax.set_ylabel("true")    for i in range(len(class_names)):        for j in range(len(class_names)):            ax.text(j, i, f"{cm[i, j]:.2f}", ha="center", fontsize=7)    buf = io.BytesIO()    fig.savefig(buf, format="png", bbox_inches="tight", dpi=110)    plt.close(fig)    buf.seek(0)    writer.add_image("confusion_matrix",                     TF.to_tensor(Image.open(buf).convert("RGB")), epoch)

Watching the confusion matrix evolve across epochs shows which class pairs the model resolves first and which it never separates. Two classes that stay confused with each other after fifty epochs are usually genuinely ambiguous in the data, and the fix is in the labels, not the architecture.

The embedding projector

Take the activations of the penultimate layer — the representation the classifier head sees — and project them to three dimensions.

Python
features, labels, images = [], [], []model.eval()with torch.no_grad():    for xb, yb in val_loader:        feats = model.backbone(xb.to(device))     # before the final layer        features.append(feats.cpu()); labels.append(yb); images.append(xb)        if len(labels) * xb.size(0) >= 1000:      # 1000 points is plenty            breakwriter.add_embedding(    torch.cat(features), metadata=torch.cat(labels).tolist(),    label_img=torch.cat(images), tag="penultimate")

Well-separated clusters, one per class, mean the network has learned a representation where classification is easy. Overlapping clouds mean it has not, and no amount of tuning the final layer will rescue it. Points of one class sitting inside another class's cluster are worth hovering over — they are frequently mislabelled examples, and finding a handful of those can be worth more than a week of architecture search.

Comparing runs

The whole point of separate run directories is overlaying them:

Text
runs/  20260518-0912_lr1e-2_bs32/  20260518-0940_lr1e-3_bs32/  20260518-1015_lr1e-4_bs32/  20260518-1102_lr1e-3_bs128/

Start TensorBoard on runs/ and all four appear on the same axes, with the directory name as the legend. For a systematic sweep, the HParams dashboard turns this into a sortable table:

Python
from torch.utils.tensorboard import SummaryWriterfor lr in [1e-2, 1e-3, 1e-4]:    for dropout in [0.0, 0.3, 0.5]:        w = SummaryWriter(f"runs/hp_lr{lr}_do{dropout}")        best_acc = train_model(lr=lr, dropout=dropout, writer=w)        w.add_hparams({"lr": lr, "dropout": dropout},                      {"hparam/best_val_acc": best_acc})        w.close()

The dashboard then gives you a parallel-coordinates view showing which hyperparameters actually correlate with the metric. Frequently the answer is "only the learning rate", which saves you from tuning the other five.

Diagnosing from the curves

ShapeDiagnosisAction
Train and val both falling togetherHealthyKeep going
Train falls, val flattens then risesOverfitting from the turning pointEarly stopping, more regularisation, more data
Both flat and high from the startUnderfitting, or the learning rate is far too lowBigger model; raise the LR by 10×
Loss spikes then recovers repeatedlyLearning rate slightly too highAdd a decay schedule
Loss jumps to nanExploding gradients or a numerical bugClip gradients; check for log of zero
Sudden drop exactly when LR is cutNormal — the optimiser can finally settleNothing; confirms the schedule works
Val loss noticeably below train lossUsually just dropout inflating the train figureNothing, if the gap is small and stable

Making it a habit rather than a chore

Instrument the run before you start it, not after it disappoints you. Re-running a six-hour job because you did not log gradient norms is a bad trade against the ninety seconds it takes to add the lines.

Encode hyperparameters in the run directory name. Six months from now, runs/exp7 tells you nothing and runs/20260518-1015_resnet18_lr1e-3_bs128_aug tells you everything. If you also write the full config as a JSON file into the same directory, the run becomes genuinely reproducible.

Keep a small fixed set of tags across all your projects — loss/train, loss/val, lr, grad_norm/total, dead_fraction/* — so that comparing an experiment from today with one from last month does not require translating between two logging schemes.

Finally, be clear about what TensorBoard is for. It is an inspection tool for the run you are looking at, not an experiment database. It has no notion of code version, dataset version or environment, and its logs get deleted when someone cleans up the disk. Once you are running dozens of experiments that matter, pair it with something that records the git commit, the config and the data hash alongside the metrics. TensorBoard answers "what is this model doing right now"; that is a different question from "what exactly produced the model we shipped in March", and you will eventually need both.