Course Content
Deep Learning with TensorFlow and PyTorch
4 sections · 15 lessons
Perceptrons and Multi-Layer Perceptrons
Imagine you are on a loan committee. Three pieces of information come across your desk for each applicant: annual income, years in current job, and number of missed payments in the last two years. You have to say yes or no.
You would not weigh those three equally. Income matters a lot. Job tenure matters somewhat. Missed payments matter a lot and in the opposite direction — more of them should push you towards no. So you invent a scoring rule: multiply each number by an importance weight, add the results, and approve the loan if the total clears some bar.
You have just described a perceptron. It was built as physical hardware in 1958 by Frank Rosenblatt, it made the front page of the New York Times, and it is the atom out of which every neural network in existence is assembled. It is also, on its own, badly broken in a way that took the field fifteen years to fully absorb. Both halves of that story are worth your time, because the fix — stacking these things in layers — is the entire architecture of a modern network.
A perceptron is a weighted vote with a threshold
Formally, a perceptron takes an input vector x=(x1,x2,…,xn), holds a weight vector w and a single bias b, and computes:
Three pieces, each with a plain-English meaning:
- Weights (wi) say how much each input counts, and in which direction. A negative weight means "more of this pushes towards no".
- Bias (b) sets how easy it is to fire at all. A large positive bias means the unit says yes unless actively talked out of it; a large negative bias means it says no unless strongly persuaded.
- Threshold function turns the continuous score into a decision.
Run our loan applicant through it. Suppose income is scaled to units of £10,000, tenure is in years, and missed payments is a count. Take w=(0.4, 0.3, −1.5) and b=−2.0. An applicant earning £60,000, three years in post, with one missed payment:
z is negative, so the answer is no — narrowly. Wipe out that single missed payment and z becomes +1.3, a comfortable yes. One binary feature flipped the decision, which is exactly the behaviour those weights encode.
1import numpy as np23def perceptron(x, w, b):4 z = np.dot(w, x) + b5 return 1 if z > 0 else 0, z67w = np.array([0.4, 0.3, -1.5])8b = -2.09print(perceptron(np.array([6.0, 3.0, 1.0]), w, b)) # (0, -0.2) reject10print(perceptron(np.array([6.0, 3.0, 0.0]), w, b)) # (1, 1.3) approveWhere the weights come from
Nobody hands you w=(0.4,0.3,−1.5). The perceptron's claim to fame is that it can find those numbers itself from labelled examples, using a rule so simple you can run it in your head:
where y is the true label, y^ is what the perceptron said, and η is a learning rate. Read the term (y−y^) carefully, because it is the whole idea:
- Prediction correct: (y−y^)=0, nothing changes. Do not fix what is not broken.
- Said 0, should have said 1: (y−y^)=+1, so weights move towards the input, making this input score higher in future.
- Said 1, should have said 0: (y−y^)=−1, so weights move away from the input.
Watch it learn the AND function from scratch. Start with w=(0,0), b=0, η=1, and cycle through the four training points.
| Step | Input | Target | z | Predicted | Error | New w | New b |
|---|---|---|---|---|---|---|---|
| 1 | (1,1) | 1 | 0 | 0 | +1 | (1,1) | 1 |
| 2 | (0,0) | 0 | 1 | 1 | −1 | (1,1) | 0 |
| 3 | (0,1) | 0 | 1 | 1 | −1 | (1,0) | −1 |
| 4 | (1,0) | 0 | 0 | 0 | 0 | (1,0) | −1 |
| 5 | (1,1) | 1 | 0 | 0 | +1 | (2,1) | 0 |
Carry on for a few more passes and it settles on w=(2,1), b=−2: the point (1,1) scores +1 and everything else scores 0 or below. It converged, and Rosenblatt proved it always will — provided a straight line separating the two classes exists at all.
That proviso is doing an enormous amount of work, and it is where the whole thing falls apart.
The problem a single perceptron cannot solve
Try XOR — "exclusive or", true when exactly one input is true.
| x1 | x2 | AND | OR | XOR |
|---|---|---|---|---|
| 0 | 0 | 0 | 0 | 0 |
| 0 | 1 | 0 | 1 | 1 |
| 1 | 0 | 0 | 1 | 1 |
| 1 | 1 | 1 | 1 | 0 |
Plot the four points on a square. XOR needs the two corners along one diagonal to be class 1, and the two corners along the other diagonal to be class 0.
x2 1 | (0,1)=1 (1,1)=0 | 0 | (0,0)=0 (1,0)=1 +-------------------------- x1 0 1 Any single straight line leaves one point on the wrong side.A perceptron's decision boundary is w1x1+w2x2+b=0, which is the equation of a straight line. You are being asked to draw one straight line with the two diagonal corners on one side and the other two on the other side. Pick up a pen and try it. It cannot be done, and the proof is a two-line contradiction: firing on (0,1) and (1,0) requires w2+b>0 and w1+b>0; adding those gives w1+w2+2b>0. Not firing on (0,0) requires b≤0, and not firing on (1,1) requires w1+w2+b≤0. Add those two and you get w1+w2+2b≤0. The same quantity cannot be both positive and non-positive.
Minsky and Papert published this in 1969, funding collapsed, and neural network research went quiet for most of the 1970s. The irony is that the fix was already obvious to the people involved: use more than one perceptron.
Stacking fixes it, and here is exactly how
A multi-layer perceptron, or MLP, arranges perceptron-like units in layers. Every unit in a layer receives every output from the layer below. The layers between input and output are called hidden, simply because nothing outside the network observes them.
Here is an explicit hand-built MLP that computes XOR. Two hidden units, one output, ReLU non-linearity (ReLU(z)=max(0,z)):
| (x1,x2) | h1=ReLU(x1+x2) | h2=ReLU(x1+x2−1) | y^=h1−2h2 | XOR target |
|---|---|---|---|---|
| (0,0) | 0 | 0 | 0 | 0 |
| (0,1) | 1 | 0 | 1 | 1 |
| (1,0) | 1 | 0 | 1 | 1 |
| (1,1) | 2 | 1 | 0 | 0 |
Look at what the hidden layer did. Unit h1 learned "at least one input is on". Unit h2 learned "both inputs are on". In that new two-dimensional space, the four original points have been moved: (0,1) and (1,0) both land at (1,0), and the two negative cases land at (0,0) and (2,1). Those are now separable by a straight line. The hidden layer did not solve XOR; it re-coordinated the problem until a single perceptron could solve it.
A hidden layer is not extra capacity bolted on. It is a change of coordinates, learned from data, chosen so that the final linear decision becomes easy.
And note again where the power comes from. If you deleted the ReLU and made both hidden units purely linear, y^ would be a linear combination of linear combinations of x1 and x2 — a straight line again, and XOR would still be impossible. The bend in the ReLU is what buys you the extra expressiveness.
Anatomy of a real MLP
Real networks are the same thing at scale. Layers are described by their width (number of units); the parameter count follows from the widths.
1import torch2import torch.nn as nn34mlp = nn.Sequential(5 nn.Linear(784, 128), # 784*128 weights + 128 biases = 100,480 parameters6 nn.ReLU(),7 nn.Linear(128, 64), # 128*64 + 64 = 8,2568 nn.ReLU(),9 nn.Linear(64, 10), # 64*10 + 10 = 65010)11total = sum(p.numel() for p in mlp.parameters())12print(f"{total:,} trainable parameters") # 109,386The same network in Keras, so you can see the mapping between the two dialects:
1from tensorflow import keras2from tensorflow.keras import layers34mlp = keras.Sequential([5 layers.Input(shape=(784,)),6 layers.Dense(128, activation="relu"),7 layers.Dense(64, activation="relu"),8 layers.Dense(10), # raw scores, no activation9])10mlp.summary() # prints the same 109,386 totalKeras calls the layer Dense and folds the activation into it as an argument; PyTorch calls it Linear and makes the activation a separate object in the stack. They compute identical things.
The output layer is decided by the task, not by taste
| Task | Output units | Final activation | Why |
|---|---|---|---|
| Binary classification | 1 | sigmoid | Squashes to a single probability in (0,1) |
| Multi-class, one label | one per class | softmax | Produces probabilities that sum to 1 across classes |
| Multi-label (tags) | one per tag | sigmoid on each | Tags are independent; they must not compete for a shared budget |
| Regression | 1 (or k) | none | The value can be any real number; squashing it would cap the range |
The classic mistake here is putting a softmax on a multi-label problem. Softmax forces the outputs to sum to 1, so a photograph that genuinely contains both a dog and a beach cannot score highly on both — raising one score mathematically requires lowering the other.
Choosing width and depth
There is no formula, but there are defensible starting points and a reliable diagnostic procedure.
| Symptom | Diagnosis | Move |
|---|---|---|
| Training loss stays high | Underfitting — model too small or training too short | Widen layers, add a layer, train longer, raise the learning rate |
| Training loss low, validation loss much higher | Overfitting — model is memorising | Shrink the model, add dropout or weight decay, get more data |
| Both losses low and close | Well-fitted | Stop fiddling |
| Loss will not move at all from step one | Broken setup, not architecture | Check the learning rate, input scaling, and that the labels line up with the inputs |
A sane default for a tabular problem is two hidden layers, widths somewhere between the input dimension and the output dimension, often tapering (128 then 64). Start deliberately too small, confirm it underfits, then grow. Starting too large hides bugs: a huge network will drive training loss to near zero even if your labels are shuffled at random, so a low training loss tells you nothing until you have seen the model fail on something.
Here is a quick sanity test worth running on any new architecture, before you spend an hour training it:
1import torch23# Take a single batch of 8 examples. A correctly wired network should be able4# to memorise 8 examples perfectly. If it cannot, the bug is in your code,5# not in your hyperparameters.6xb, yb = next(iter(train_loader))7xb, yb = xb[:8], yb[:8]89opt = torch.optim.Adam(mlp.parameters(), lr=1e-3)10loss_fn = torch.nn.CrossEntropyLoss()11for step in range(300):12 opt.zero_grad()13 loss = loss_fn(mlp(xb), yb)14 loss.backward()15 opt.step()16print(loss.item()) # should be essentially 0.0; if it plateaus, debug nowCarrying this into real work
Three habits follow from what a perceptron can and cannot do.
When a model refuses to learn, ask whether the task is linearly separable in the features you gave it. A logistic regression is a perceptron with a smoother threshold, and it fails on XOR-shaped structure for exactly the reasons shown above. If your features are "user age" and "time of day" and the true pattern is "young people at night or older people in the morning", no linear model will find it. You either add a hidden layer or engineer an interaction feature by hand.
Treat the hidden layers as a learned representation, not as magic. Their job is to bend the space until the last layer's straight-line decision works. This is why inspecting hidden activations is a genuinely useful debugging move: if every unit in a layer outputs zero for every input, that layer has died and is contributing nothing.
Do not skip the non-linearity. It is a one-word omission that silently reduces a twelve-layer network to a linear model, and it will not raise an error. It will just train to a mediocre score and leave you wondering why depth did not help.