Course Content
Deep Learning Essentials
13 sections · 61 lessons
What is the difference between Single-Layer and Multi-Layer Perceptron?
What you need to know
The single-layer perceptron
Frank Rosenblatt's perceptron (1958) computes a weighted sum and a step function:
output = 1 if (w1*x1 + w2*x2 + b) > 0 else 0The set of points where w1*x1 + w2*x2 + b = 0 is a straight line. Everything on one side is class 1, everything on the other is class 0. So the perceptron can only solve problems where one straight line separates the classes. AND and OR are like that. XOR is not.
x1 x2 XOR0 0 00 1 11 0 11 1 0The 1s sit on one diagonal and the 0s on the other. Any line that separates (0,1) and (1,0) from (0,0) also cuts off (1,1) wrongly. In 1969 Minsky and Papert highlighted this limit, and interest in neural networks dropped for years.
The multi-layer perceptron
Add one hidden layer with ReLU and XOR becomes easy. These hand-picked weights solve it:
h1 = ReLU(x1 + x2)h2 = ReLU(x1 + x2 - 1)out = h1 - 2*h2(0,0): h1=0, h2=0 -> 0(0,1): h1=1, h2=0 -> 1(1,0): h1=1, h2=0 -> 1(1,1): h1=2, h2=1 -> 0The hidden layer maps the four points into a new space where the classes can be separated. h2 only turns on when both inputs are 1, and subtracting it cancels the output. In a real MLP, backpropagation finds weights like these on its own.
Single-layer perceptron
- Inputs connect straight to output
- One straight-line boundary
- Cannot learn XOR
- Trained by the perceptron rule
Multi-layer perceptron
- One or more hidden layers
- Curved, complex boundaries
- Learns XOR and far more
- Trained by backpropagation
Why training differs
The perceptron rule updates weights only when a prediction is wrong: w = w + lr * (y − prediction) * x. It works because there is only one layer, so the error directly tells each weight what to do. In an MLP, the hidden neurons have no target of their own. Backpropagation uses the chain rule to pass the output error back and give each hidden weight its share of the blame. That also needs a differentiable activation, which is why MLPs use sigmoid, tanh or ReLU instead of a step function.
A real-life example
A bank's first fraud model was a linear score on two features: transaction amount and distance from the cardholder's home. But their customers' normal behaviour looked like this: small purchases near home (groceries, fuel) and large purchases far from home (flights, hotels on holiday). The fraud they kept missing sat in the other two corners: small test charges far away, and large purchases near home with a cloned card. That is the XOR shape. No single line can put both fraud corners on one side and both normal corners on the other.
An MLP with one hidden layer of 16 neurons learned both corners, and fraud recall on that pattern rose from 0.31 to 0.77. The hidden neurons had learned sub-patterns, "small and far" and "large and near", that the output neuron could combine.
Follow-up questions to expect
- "Can a single perceptron with a sigmoid learn XOR?" — No. A sigmoid still gives a single linear boundary at 0.5. Logistic regression is also a linear classifier.
- "How many hidden neurons are needed for XOR?" — Two are enough, as in the example above.
- "Does the perceptron rule always converge?" — Only if the data is linearly separable. Otherwise it keeps changing forever.