Deep Learning Essentials

Course Content

Deep Learning Essentials

13 sections · 61 lessons

What is the difference between Single-Layer and Multi-Layer Perceptron?


A two-neuron hidden layer that solves XORInputs: (0,0) (0,1) (1,0) (1,1)h1 = ReLU(x1 + x2): 0, 1, 1, 2h2 = ReLU(x1 + x2 − 1): 0, 0, 0, 1out = h1 − 2·h2: 0, 1, 1, 0
h2 fires only when both inputs are on, and subtracting it bends the boundary no single perceptron line could draw.

What you need to know

The single-layer perceptron

Frank Rosenblatt's perceptron (1958) computes a weighted sum and a step function:

Text
output = 1 if (w1*x1 + w2*x2 + b) > 0 else 0

The set of points where w1*x1 + w2*x2 + b = 0 is a straight line. Everything on one side is class 1, everything on the other is class 0. So the perceptron can only solve problems where one straight line separates the classes. AND and OR are like that. XOR is not.

Text
x1  x2  XOR0   0   00   1   11   0   11   1   0

The 1s sit on one diagonal and the 0s on the other. Any line that separates (0,1) and (1,0) from (0,0) also cuts off (1,1) wrongly. In 1969 Minsky and Papert highlighted this limit, and interest in neural networks dropped for years.

The multi-layer perceptron

Add one hidden layer with ReLU and XOR becomes easy. These hand-picked weights solve it:

Text
h1  = ReLU(x1 + x2)h2  = ReLU(x1 + x2 - 1)out = h1 - 2*h2(0,0): h1=0, h2=0 -> 0(0,1): h1=1, h2=0 -> 1(1,0): h1=1, h2=0 -> 1(1,1): h1=2, h2=1 -> 0

The hidden layer maps the four points into a new space where the classes can be separated. h2 only turns on when both inputs are 1, and subtracting it cancels the output. In a real MLP, backpropagation finds weights like these on its own.

Single-layer perceptron

  • Inputs connect straight to output
  • One straight-line boundary
  • Cannot learn XOR
  • Trained by the perceptron rule

Multi-layer perceptron

  • One or more hidden layers
  • Curved, complex boundaries
  • Learns XOR and far more
  • Trained by backpropagation

Why training differs

The perceptron rule updates weights only when a prediction is wrong: w = w + lr * (y − prediction) * x. It works because there is only one layer, so the error directly tells each weight what to do. In an MLP, the hidden neurons have no target of their own. Backpropagation uses the chain rule to pass the output error back and give each hidden weight its share of the blame. That also needs a differentiable activation, which is why MLPs use sigmoid, tanh or ReLU instead of a step function.

A real-life example

A bank's first fraud model was a linear score on two features: transaction amount and distance from the cardholder's home. But their customers' normal behaviour looked like this: small purchases near home (groceries, fuel) and large purchases far from home (flights, hotels on holiday). The fraud they kept missing sat in the other two corners: small test charges far away, and large purchases near home with a cloned card. That is the XOR shape. No single line can put both fraud corners on one side and both normal corners on the other.

An MLP with one hidden layer of 16 neurons learned both corners, and fraud recall on that pattern rose from 0.31 to 0.77. The hidden neurons had learned sub-patterns, "small and far" and "large and near", that the output neuron could combine.

Follow-up questions to expect

  • "Can a single perceptron with a sigmoid learn XOR?" — No. A sigmoid still gives a single linear boundary at 0.5. Logistic regression is also a linear classifier.
  • "How many hidden neurons are needed for XOR?" — Two are enough, as in the example above.
  • "Does the perceptron rule always converge?" — Only if the data is linearly separable. Otherwise it keeps changing forever.