Course Content
Deep Learning with TensorFlow and PyTorch
4 sections · 15 lessons
Custom Neural Networks in TensorFlow and PyTorch
You are predicting how long a delivery will take. You have two very different kinds of input: a satellite image of the route, and a handful of numbers — distance, time of day, driver rating, parcel weight. The image needs convolutional layers. The numbers need a small dense stack. Somewhere near the end the two have to meet and produce a single prediction.
Now try to express that as Sequential.
You cannot. Sequential means exactly one thing: layer 1 feeds layer 2 feeds layer 3, one input at the top, one output at the bottom, no branches. Your architecture has two inputs that travel separately and then merge. It is a graph, not a stack, and the stack API has no vocabulary for it.
The same wall appears for a residual connection — where a layer's input is added to its output, so the network can learn a small correction rather than a whole transformation. That is a wire that skips forward past a layer, and again, a stack cannot say it.
This is the moment when everyone graduates from the beginner API. What follows is how to say things that Sequential cannot, in both frameworks.
An architecture is a directed graph
Before writing code, sketch the data flow. Boxes for operations, arrows for tensors, and a shape annotation on every arrow.
image (128,128,3) metadata (12,) | | Conv+Pool Dense 32 | | Conv+Pool Dense 16 | | Flatten (2048,) | | | +----------- concat -------+ | (2064,) Dense 64 | Dense 1 -> predicted minutesDoing this on paper first is not ceremony. Most architecture bugs are shape mismatches, and shape mismatches are visible in the sketch before they are visible in a stack trace. The concatenation above works because both branches have been reduced to rank-1 tensors per example; if the image branch still had spatial dimensions, the concat would fail and the sketch would have shown you why.
Keras: three APIs, in increasing order of freedom
Sequential — for straight stacks only
1from tensorflow import keras2from tensorflow.keras import layers34model = keras.Sequential([5 layers.Input(shape=(784,)),6 layers.Dense(128, activation="relu"),7 layers.Dropout(0.3),8 layers.Dense(64, activation="relu"),9 layers.Dense(10),10])Readable, concise, and correct for perhaps 60% of real models. Use it when it fits and move on.
Functional — for anything with a shape
The Functional API treats layers as callable objects. You call a layer on a tensor and get a tensor back, so the code you write is the graph.
1from tensorflow import keras2from tensorflow.keras import layers34image_in = keras.Input(shape=(128, 128, 3), name="image")5meta_in = keras.Input(shape=(12,), name="metadata")67# Image branch8x = layers.Conv2D(32, 3, activation="relu")(image_in)9x = layers.MaxPooling2D()(x)10x = layers.Conv2D(64, 3, activation="relu")(x)11x = layers.GlobalAveragePooling2D()(x) # (batch, 64)1213# Metadata branch14m = layers.Dense(32, activation="relu")(meta_in)15m = layers.Dense(16, activation="relu")(m) # (batch, 16)1617# Merge and predict18merged = layers.Concatenate()([x, m]) # (batch, 80)19h = layers.Dense(64, activation="relu")(merged)20out = layers.Dense(1, name="minutes")(h)2122model = keras.Model(inputs=[image_in, meta_in], outputs=out)A residual block is equally direct — you keep a reference to the tensor and add it back later:
1def residual_block(x, units):2 shortcut = x # remember the input3 y = layers.Dense(units, activation="relu")(x)4 y = layers.Dense(units)(y) # no activation yet5 y = layers.Add()([y, shortcut]) # the skip connection6 return layers.Activation("relu")(y) # activate after addingNote the ordering: activate after the addition, not before. Putting the ReLU inside the branch and then adding means the shortcut path passes through a non-linearity, which defeats the purpose — the whole point is an unobstructed route for the gradient to travel backwards.
Note also that Add requires both tensors to have identical shapes. If the block changes the number of units, the shortcut needs a projection: shortcut = layers.Dense(units)(x). Forgetting this is the most common error when writing residual networks by hand.
Subclassing — when the forward pass has logic in it
Some models need control flow: a loop whose length depends on the input, a branch taken only during training, a computation that cannot be expressed as a static graph. For those, subclass keras.Model and write call() as ordinary Python.
1class GatedRegressor(keras.Model):2 def __init__(self, units=64, **kwargs):3 super().__init__(**kwargs)4 self.hidden = layers.Dense(units, activation="relu")5 self.gate = layers.Dense(units, activation="sigmoid")6 self.out = layers.Dense(1)78 def call(self, inputs, training=False):9 h = self.hidden(inputs)10 g = self.gate(inputs)11 h = h * g # learned per-unit gating12 if training:13 h = tf.nn.dropout(h, rate=0.2) # only during training14 return self.out(h)| API | Can express | Gives you | Costs you |
|---|---|---|---|
| Sequential | Linear stacks | Shortest code, full summary(), easy save/load | No branches, one input, one output |
| Functional | Any static graph | Shape checking at build time, plottable graph, easy save/load | No data-dependent control flow |
| Subclassing | Anything Python can express | Total freedom | No shape checks until you call it; summary() is uninformative until built; saving needs get_config() |
The practical advice is to use the least powerful API that expresses your model. Functional catches shape errors when you build the graph, which is seconds after you write the bug. Subclassing catches them on the first forward pass, which may be after a long data-loading step.
PyTorch: one pattern for everything
PyTorch has no equivalent split. Everything is a subclass of nn.Module with two methods: __init__ creates the layers, forward defines what happens to the data. Because forward is plain Python, branches, loops, and conditionals all work without ceremony.
1import torch2import torch.nn as nn34class DeliveryTimeModel(nn.Module):5 def __init__(self, n_meta=12):6 super().__init__() # never forget this line7 self.conv = nn.Sequential(8 nn.Conv2d(3, 32, 3, padding=1), nn.ReLU(), nn.MaxPool2d(2),9 nn.Conv2d(32, 64, 3, padding=1), nn.ReLU(),10 nn.AdaptiveAvgPool2d(1), nn.Flatten(), # -> (batch, 64)11 )12 self.meta = nn.Sequential(13 nn.Linear(n_meta, 32), nn.ReLU(),14 nn.Linear(32, 16), nn.ReLU(), # -> (batch, 16)15 )16 self.head = nn.Sequential(17 nn.Linear(64 + 16, 64), nn.ReLU(),18 nn.Linear(64, 1),19 )2021 def forward(self, image, metadata):22 a = self.conv(image)23 b = self.meta(metadata)24 return self.head(torch.cat([a, b], dim=1)) # dim=1 is the feature axis2526model = DeliveryTimeModel()27out = model(torch.randn(8, 3, 128, 128), torch.randn(8, 12))28print(out.shape) # torch.Size([8, 1])Two details in there matter. super().__init__() is what sets up the machinery that tracks parameters; omit it and you get AttributeError: cannot assign module before Module.__init__() call as soon as you assign the first layer. And torch.cat(..., dim=1) concatenates along the feature axis — dim=0 would stack the two branches as extra examples, which silently produces a batch of 16 half-filled rows instead of 8 complete ones.
A residual block, same idea:
1class ResidualBlock(nn.Module):2 def __init__(self, dim):3 super().__init__()4 self.net = nn.Sequential(5 nn.Linear(dim, dim), nn.ReLU(),6 nn.Linear(dim, dim),7 )89 def forward(self, x):10 return torch.relu(x + self.net(x)) # add first, then activateThe silent bug: plain lists do not register parameters
Suppose you want a variable number of layers. The obvious code is wrong:
1class Broken(nn.Module):2 def __init__(self, depth=4, dim=64):3 super().__init__()4 self.layers = [nn.Linear(dim, dim) for _ in range(depth)] # BUG56 def forward(self, x):7 for layer in self.layers:8 x = torch.relu(layer(x))9 return x1011print(sum(p.numel() for p in Broken().parameters())) # 0Zero parameters. The forward pass runs perfectly and produces sensible-looking numbers, so nothing errors. But model.parameters() is empty, so the optimiser has nothing to update, so the model never learns — and .to(device) will not move those layers either, giving you a device mismatch on the first GPU run.
nn.Module discovers submodules by inspecting attributes assigned directly on self. A Python list is one attribute containing objects it does not look inside. The fix is a container that nn.Module understands:
1class Fixed(nn.Module):2 def __init__(self, depth=4, dim=64):3 super().__init__()4 self.layers = nn.ModuleList([nn.Linear(dim, dim) for _ in range(depth)])56 def forward(self, x):7 for layer in self.layers:8 x = torch.relu(layer(x))9 return x1011print(sum(p.numel() for p in Fixed().parameters())) # 16,640Print your parameter count immediately after building any model. A count of zero, or a count that is wildly lower than you expect, is the cheapest bug detector you have.
The related containers are nn.ModuleDict for name-keyed submodules, and nn.ParameterList for bare tensors you want trained. Note also that nn.Sequential is itself just an nn.Module that calls its children in order — which is why nesting nn.Sequential inside a custom module, as in the delivery model above, works exactly as you would hope.
Weight sharing, which comes free
Sometimes two parts of a model should use the same weights — a siamese network comparing two images, or a decoder that ties its output projection to the input embedding. Both frameworks handle this by simply reusing the object.
1# PyTorch: call the same module twice2class Siamese(nn.Module):3 def __init__(self):4 super().__init__()5 self.encoder = nn.Sequential(nn.Linear(128, 64), nn.ReLU(),6 nn.Linear(64, 32))78 def forward(self, a, b):9 ea, eb = self.encoder(a), self.encoder(b) # same weights, twice10 return torch.nn.functional.cosine_similarity(ea, eb, dim=1)1# Keras Functional: call the same layer object on two tensors2shared = keras.Sequential([layers.Dense(64, activation="relu"),3 layers.Dense(32)])4in_a, in_b = keras.Input((128,)), keras.Input((128,))5sim = layers.Dot(axes=1, normalize=True)([shared(in_a), shared(in_b)])6model = keras.Model([in_a, in_b], sim)The parameter count confirms it: a siamese encoder appears once in model.parameters(), not twice. Gradients from both calls accumulate into the same tensors, which is precisely the intent.
Custom layers
Occasionally you need an operation no built-in layer provides. Both frameworks make this a small amount of code.
1class ScaledLinear(nn.Module):2 """Linear layer whose output is multiplied by a learned per-unit scale."""3 def __init__(self, in_features, out_features):4 super().__init__()5 self.weight = nn.Parameter(torch.randn(out_features, in_features) * 0.02)6 self.bias = nn.Parameter(torch.zeros(out_features))7 self.scale = nn.Parameter(torch.ones(out_features))89 def forward(self, x):10 return (x @ self.weight.T + self.bias) * self.scalenn.Parameter is the key. It is a tensor tagged as "this is trainable"; assign it as an attribute and it appears in parameters() automatically. A plain torch.randn assigned to self is treated as a constant and never trained — the same class of silent failure as the plain list.
1class ScaledDense(layers.Layer):2 def __init__(self, units):3 super().__init__()4 self.units = units56 def build(self, input_shape): # called on first use, once shapes are known7 self.w = self.add_weight(shape=(input_shape[-1], self.units),8 initializer="glorot_uniform", trainable=True)9 self.b = self.add_weight(shape=(self.units,),10 initializer="zeros", trainable=True)11 self.s = self.add_weight(shape=(self.units,),12 initializer="ones", trainable=True)1314 def call(self, x):15 return (tf.matmul(x, self.w) + self.b) * self.sKeras splits construction into __init__ (things you know immediately) and build (things that depend on the input shape). That is why you can write Dense(64) without saying how many inputs it takes — the shape is inferred the first time data arrives.
Inspecting what you built
Never assume the model you wrote is the model you meant. Check.
model.summary() # Keras: layer names, output shapes, param countskeras.utils.plot_model(model, show_shapes=True, to_file="arch.png")1# PyTorch: parameter counts, by layer and in total2for name, p in model.named_parameters():3 print(f"{name:30s} {str(tuple(p.shape)):20s} {p.numel():>10,}")4total = sum(p.numel() for p in model.parameters() if p.requires_grad)5print(f"{'TOTAL':30s} {'':20s} {total:>10,}")67# Shape-trace the forward pass with a hook on every layer8def trace(name):9 return lambda m, i, o: print(f"{name:25s} -> {tuple(o.shape)}")10for name, layer in model.named_children():11 layer.register_forward_hook(trace(name))12model(torch.randn(2, 3, 128, 128), torch.randn(2, 12))That hook trick is the fastest way to find a shape bug. Rather than reading the error and guessing which layer produced the wrong tensor, you get a printed shape after every step and can see exactly where it went wrong.
Choosing an approach when you sit down to build
Start from the sketch, and let the shape of the graph pick the API. If the data flows straight through, use Sequential in Keras — there is no prize for writing more code than the problem needs. If there are branches, merges or skips but the graph is fixed, use the Functional API, because building the graph up front catches shape errors at the point you write them. If the forward pass contains a decision that depends on the data — a loop over a variable-length sequence, a branch on a flag — subclass. In PyTorch the question does not arise; you subclass nn.Module for everything, and the only real discipline is remembering that submodules and parameters must be assigned as attributes, or wrapped in ModuleList, to be seen at all.
Whichever route you take, do two things before you start training. Run one forward pass on a batch of two random tensors of the right shape, and confirm the output shape is what you intended. Then print the total trainable parameter count and check it against a rough hand estimate. Those two checks take fifteen seconds and catch the great majority of architecture bugs — including the ones that would otherwise train quietly for an hour and produce a model that never learned anything at all.