Deep Learning with TensorFlow and PyTorch

Perceptrons and Multi-Layer Perceptrons


Imagine you are on a loan committee. Three pieces of information come across your desk for each applicant: annual income, years in current job, and number of missed payments in the last two years. You have to say yes or no.

You would not weigh those three equally. Income matters a lot. Job tenure matters somewhat. Missed payments matter a lot and in the opposite direction — more of them should push you towards no. So you invent a scoring rule: multiply each number by an importance weight, add the results, and approve the loan if the total clears some bar.

You have just described a perceptron. It was built as physical hardware in 1958 by Frank Rosenblatt, it made the front page of the New York Times, and it is the atom out of which every neural network in existence is assembled. It is also, on its own, badly broken in a way that took the field fifteen years to fully absorb. Both halves of that story are worth your time, because the fix — stacking these things in layers — is the entire architecture of a modern network.

The XOR table no single straight line can split0110x2 = 0x2 = 1x1 = 0x1 = 1The two filled cells are class 1 and sit on a diagonal; the two class-0 cells sit on the other.
One perceptron draws one line, and no line separates a diagonal — the hidden layer is what bends it.

A perceptron is a weighted vote with a threshold

Formally, a perceptron takes an input vector x=(x1,x2,…,xn)x = (x_1, x_2, \dots, x_n), holds a weight vector ww and a single bias bb, and computes:

z=∑i=1nwixi+b,y^={1if z>00otherwisez = \sum_{i=1}^{n} w_i x_i + b, \qquad \hat{y} = \begin{cases} 1 & \text{if } z > 0 \\ 0 & \text{otherwise} \end{cases}

Three pieces, each with a plain-English meaning:

  • Weights (wiw_i) say how much each input counts, and in which direction. A negative weight means "more of this pushes towards no".
  • Bias (bb) sets how easy it is to fire at all. A large positive bias means the unit says yes unless actively talked out of it; a large negative bias means it says no unless strongly persuaded.
  • Threshold function turns the continuous score into a decision.

Run our loan applicant through it. Suppose income is scaled to units of £10,000, tenure is in years, and missed payments is a count. Take w=(0.4, 0.3, −1.5)w = (0.4,\ 0.3,\ -1.5) and b=−2.0b = -2.0. An applicant earning £60,000, three years in post, with one missed payment:

z=0.4(6)+0.3(3)+(−1.5)(1)−2.0=2.4+0.9−1.5−2.0=−0.2z = 0.4(6) + 0.3(3) + (-1.5)(1) - 2.0 = 2.4 + 0.9 - 1.5 - 2.0 = -0.2

zz is negative, so the answer is no — narrowly. Wipe out that single missed payment and zz becomes +1.3+1.3, a comfortable yes. One binary feature flipped the decision, which is exactly the behaviour those weights encode.

Python
import numpy as npdef perceptron(x, w, b):    z = np.dot(w, x) + b    return 1 if z > 0 else 0, zw = np.array([0.4, 0.3, -1.5])b = -2.0print(perceptron(np.array([6.0, 3.0, 1.0]), w, b))   # (0, -0.2)  rejectprint(perceptron(np.array([6.0, 3.0, 0.0]), w, b))   # (1,  1.3)  approve

Where the weights come from

Nobody hands you w=(0.4,0.3,−1.5)w = (0.4, 0.3, -1.5). The perceptron's claim to fame is that it can find those numbers itself from labelled examples, using a rule so simple you can run it in your head:

w←w+η (y−y^) x,b←b+η (y−y^)w \leftarrow w + \eta\,(y - \hat{y})\,x, \qquad b \leftarrow b + \eta\,(y - \hat{y})

where yy is the true label, y^\hat{y} is what the perceptron said, and η\eta is a learning rate. Read the term (y−y^)(y - \hat{y}) carefully, because it is the whole idea:

  • Prediction correct: (y−y^)=0(y - \hat{y}) = 0, nothing changes. Do not fix what is not broken.
  • Said 0, should have said 1: (y−y^)=+1(y - \hat{y}) = +1, so weights move towards the input, making this input score higher in future.
  • Said 1, should have said 0: (y−y^)=−1(y - \hat{y}) = -1, so weights move away from the input.

Watch it learn the AND function from scratch. Start with w=(0,0)w = (0, 0), b=0b = 0, η=1\eta = 1, and cycle through the four training points.

StepInputTargetzzPredictedErrorNew wwNew bb
1(1,1)100+1(1,1)1
2(0,0)011−1(1,1)0
3(0,1)011−1(1,0)−1
4(1,0)0000(1,0)−1
5(1,1)100+1(2,1)0

Carry on for a few more passes and it settles on w=(2,1)w = (2, 1), b=−2b = -2: the point (1,1)(1,1) scores +1+1 and everything else scores 00 or below. It converged, and Rosenblatt proved it always will — provided a straight line separating the two classes exists at all.

That proviso is doing an enormous amount of work, and it is where the whole thing falls apart.

The problem a single perceptron cannot solve

Try XOR — "exclusive or", true when exactly one input is true.

x1x_1x2x_2ANDORXOR
00000
01011
10011
11110

Plot the four points on a square. XOR needs the two corners along one diagonal to be class 1, and the two corners along the other diagonal to be class 0.

Text
   x2    1 |  (0,1)=1        (1,1)=0      |    0 |  (0,0)=0        (1,0)=1      +-------------------------- x1         0               1  Any single straight line leaves one point on the wrong side.

A perceptron's decision boundary is w1x1+w2x2+b=0w_1x_1 + w_2x_2 + b = 0, which is the equation of a straight line. You are being asked to draw one straight line with the two diagonal corners on one side and the other two on the other side. Pick up a pen and try it. It cannot be done, and the proof is a two-line contradiction: firing on (0,1)(0,1) and (1,0)(1,0) requires w2+b>0w_2 + b > 0 and w1+b>0w_1 + b > 0; adding those gives w1+w2+2b>0w_1 + w_2 + 2b > 0. Not firing on (0,0)(0,0) requires b≤0b \le 0, and not firing on (1,1)(1,1) requires w1+w2+b≤0w_1 + w_2 + b \le 0. Add those two and you get w1+w2+2b≤0w_1 + w_2 + 2b \le 0. The same quantity cannot be both positive and non-positive.

Minsky and Papert published this in 1969, funding collapsed, and neural network research went quiet for most of the 1970s. The irony is that the fix was already obvious to the people involved: use more than one perceptron.

Stacking fixes it, and here is exactly how

A multi-layer perceptron, or MLP, arranges perceptron-like units in layers. Every unit in a layer receives every output from the layer below. The layers between input and output are called hidden, simply because nothing outside the network observes them.

Here is an explicit hand-built MLP that computes XOR. Two hidden units, one output, ReLU non-linearity (ReLU(z)=max⁡(0,z)\text{ReLU}(z) = \max(0, z)):

h1=ReLU(x1+x2+0),h2=ReLU(x1+x2−1),y^=h1−2h2h_1 = \text{ReLU}(x_1 + x_2 + 0), \qquad h_2 = \text{ReLU}(x_1 + x_2 - 1), \qquad \hat{y} = h_1 - 2h_2

(x1,x2)(x_1,x_2)h1=ReLU(x1+x2)h_1 = \text{ReLU}(x_1{+}x_2)h2=ReLU(x1+x2−1)h_2 = \text{ReLU}(x_1{+}x_2{-}1)y^=h1−2h2\hat{y} = h_1 - 2h_2XOR target
(0,0)0000
(0,1)1011
(1,0)1011
(1,1)2100

Look at what the hidden layer did. Unit h1h_1 learned "at least one input is on". Unit h2h_2 learned "both inputs are on". In that new two-dimensional space, the four original points have been moved: (0,1)(0,1) and (1,0)(1,0) both land at (1,0)(1,0), and the two negative cases land at (0,0)(0,0) and (2,1)(2,1). Those are now separable by a straight line. The hidden layer did not solve XOR; it re-coordinated the problem until a single perceptron could solve it.

A hidden layer is not extra capacity bolted on. It is a change of coordinates, learned from data, chosen so that the final linear decision becomes easy.

And note again where the power comes from. If you deleted the ReLU and made both hidden units purely linear, y^\hat{y} would be a linear combination of linear combinations of x1x_1 and x2x_2 — a straight line again, and XOR would still be impossible. The bend in the ReLU is what buys you the extra expressiveness.

Anatomy of a real MLP

Real networks are the same thing at scale. Layers are described by their width (number of units); the parameter count follows from the widths.

Python
import torchimport torch.nn as nnmlp = nn.Sequential(    nn.Linear(784, 128),   # 784*128 weights + 128 biases = 100,480 parameters    nn.ReLU(),    nn.Linear(128, 64),    # 128*64 + 64                  =   8,256    nn.ReLU(),    nn.Linear(64, 10),     # 64*10 + 10                   =     650)total = sum(p.numel() for p in mlp.parameters())print(f"{total:,} trainable parameters")   # 109,386

The same network in Keras, so you can see the mapping between the two dialects:

Python
from tensorflow import kerasfrom tensorflow.keras import layersmlp = keras.Sequential([    layers.Input(shape=(784,)),    layers.Dense(128, activation="relu"),    layers.Dense(64,  activation="relu"),    layers.Dense(10),                       # raw scores, no activation])mlp.summary()   # prints the same 109,386 total

Keras calls the layer Dense and folds the activation into it as an argument; PyTorch calls it Linear and makes the activation a separate object in the stack. They compute identical things.

The output layer is decided by the task, not by taste

TaskOutput unitsFinal activationWhy
Binary classification1sigmoidSquashes to a single probability in (0,1)
Multi-class, one labelone per classsoftmaxProduces probabilities that sum to 1 across classes
Multi-label (tags)one per tagsigmoid on eachTags are independent; they must not compete for a shared budget
Regression1 (or kk)noneThe value can be any real number; squashing it would cap the range

The classic mistake here is putting a softmax on a multi-label problem. Softmax forces the outputs to sum to 1, so a photograph that genuinely contains both a dog and a beach cannot score highly on both — raising one score mathematically requires lowering the other.

Choosing width and depth

There is no formula, but there are defensible starting points and a reliable diagnostic procedure.

SymptomDiagnosisMove
Training loss stays highUnderfitting — model too small or training too shortWiden layers, add a layer, train longer, raise the learning rate
Training loss low, validation loss much higherOverfitting — model is memorisingShrink the model, add dropout or weight decay, get more data
Both losses low and closeWell-fittedStop fiddling
Loss will not move at all from step oneBroken setup, not architectureCheck the learning rate, input scaling, and that the labels line up with the inputs

A sane default for a tabular problem is two hidden layers, widths somewhere between the input dimension and the output dimension, often tapering (128 then 64). Start deliberately too small, confirm it underfits, then grow. Starting too large hides bugs: a huge network will drive training loss to near zero even if your labels are shuffled at random, so a low training loss tells you nothing until you have seen the model fail on something.

Here is a quick sanity test worth running on any new architecture, before you spend an hour training it:

Python
import torch# Take a single batch of 8 examples. A correctly wired network should be able# to memorise 8 examples perfectly. If it cannot, the bug is in your code,# not in your hyperparameters.xb, yb = next(iter(train_loader))xb, yb = xb[:8], yb[:8]opt = torch.optim.Adam(mlp.parameters(), lr=1e-3)loss_fn = torch.nn.CrossEntropyLoss()for step in range(300):    opt.zero_grad()    loss = loss_fn(mlp(xb), yb)    loss.backward()    opt.step()print(loss.item())   # should be essentially 0.0; if it plateaus, debug now

Carrying this into real work

Three habits follow from what a perceptron can and cannot do.

When a model refuses to learn, ask whether the task is linearly separable in the features you gave it. A logistic regression is a perceptron with a smoother threshold, and it fails on XOR-shaped structure for exactly the reasons shown above. If your features are "user age" and "time of day" and the true pattern is "young people at night or older people in the morning", no linear model will find it. You either add a hidden layer or engineer an interaction feature by hand.

Treat the hidden layers as a learned representation, not as magic. Their job is to bend the space until the last layer's straight-line decision works. This is why inspecting hidden activations is a genuinely useful debugging move: if every unit in a layer outputs zero for every input, that layer has died and is contributing nothing.

Do not skip the non-linearity. It is a one-word omission that silently reduces a twelve-layer network to a linear model, and it will not raise an error. It will just train to a mediocre score and leave you wondering why depth did not help.