Deep Learning Essentials

Course Content

Deep Learning Essentials

13 sections · 61 lessons

Explain the Boltzmann machine. What is a Restricted Boltzmann machine?


What you need to know

The Boltzmann machine

  • Units are binary (0 or 1) and stochastic: each turns on with a probability, not deterministically.
  • Units are split into visible (the data) and hidden (learned features).
  • In a general Boltzmann machine, every unit connects to every other, including hidden to hidden.
  • Training adjusts weights so that real data gets low energy. But computing the needed statistics requires long random sampling over the whole network, which is far too slow for real use.

The restriction

An RBM keeps only connections between the visible layer and the hidden layer. No visible–visible, no hidden–hidden. The graph is bipartite.

Text
energy(v, h) = - sum(a_i * v_i) - sum(b_j * h_j) - sum(v_i * W_ij * h_j)P(h_j = 1 | v) = sigmoid(b_j + sum_i v_i * W_ij)P(v_i = 1 | h) = sigmoid(a_i + sum_j W_ij * h_j)

Because hidden units do not talk to each other, all of them can be sampled at once given the visible layer, and vice versa. That makes training fast enough.

Contrastive divergence (CD-1), in words

  1. Put a training example on the visible layer and sample the hidden layer.
  2. From that hidden sample, reconstruct the visible layer, then sample the hidden layer again.
  3. Update each weight by (data correlation of v_i h_j) minus (reconstruction correlation of v_i h_j), times a learning rate.

This pushes the model to make real data more likely than its own reconstructions.

Why it matters historically

  • In 2006, Hinton and colleagues stacked RBMs into deep belief networks, training one layer at a time. This was one of the first ways to train deep networks well, and helped start the deep learning revival.
  • RBMs were part of strong collaborative-filtering systems in the Netflix Prize competition.
  • In 2024, the Nobel Prize in Physics went to John Hopfield and Geoffrey Hinton; the Boltzmann machine was one of the cited contributions.

They faded because ReLU, better initialisation, batch normalisation and large labelled datasets made layer-wise pretraining unnecessary, and VAEs, GANs and diffusion models became better generative models.

A real-life example

Around 2007, a movie site wants to recommend films. Each user is a visible layer of 17,000 binary units: 1 if they liked a movie, 0 otherwise. The RBM has 100 hidden units.

After training, hidden units act like tastes: one turns on for users who like 1990s action films, another for animated family films. To recommend, you put a user's likes on the visible layer, compute the hidden units' probabilities, then reconstruct the visible layer. Movies the user has not rated but that get a high reconstructed probability become recommendations.

A modern team would solve the same problem with matrix factorisation or a two-tower embedding model, which train faster and scale better. That comparison is a good way to end your interview answer.

Follow-up questions to expect

  • "Why 'restricted'?" — Because connections inside each layer are removed; only visible-to-hidden connections remain.
  • "What is a deep belief network?" — A stack of RBMs trained one layer at a time, each on the hidden activations of the one below, often fine-tuned afterwards with backpropagation.
  • "Would you use an RBM today?" — Rarely. I would use an autoencoder or VAE for representation learning and diffusion models or GANs for generation.