Deep Learning with TensorFlow and PyTorch

What Deep Learning Is and Why It Matters


Here is a job. A postal service hands you 60,000 scanned images of handwritten digits, each one a 28×28 grid of grey pixels, and asks you to write software that reads them. Get it wrong and letters go to the wrong city.

You are a competent programmer, so you start where competent programmers start: you look at the data and you think about features. A 7 has a long horizontal stroke at the top. An 8 has two enclosed loops. A 1 is mostly a single vertical line, so the ratio of ink-height to ink-width should be large. You write a function that counts enclosed regions. You write another that measures stroke thickness. You write a third that finds the centre of mass of the ink.

Then you look at the actual handwriting. Someone's 7 has a crossbar; someone else's does not. Someone's 4 is closed at the top and looks like a 9. Someone writes a 2 with a loop at the bottom, and now your enclosed-region counter says it is a 6. A 1 written at a slant has the same height-to-width ratio as a 7. You add exceptions. The exceptions collide with each other. After three weeks your accuracy has crawled to about 92%, which sounds respectable until you work out that it means roughly one letter in twelve gets misrouted.

The problem is not that you are bad at this. The problem is that you are trying to write down knowledge you do not consciously possess. You recognise a handwritten 7 in about 200 milliseconds and you have no idea how. The rules are not in your head as rules. They are in your head as something else, and the entire field of deep learning exists because someone worked out how to build that something else out of arithmetic.

What the 28 by 28 digit becomes, layer by layer784 raw grey pixel valuesEdges and strokes at particular anglesLoops, junctions, line endingsWhole-digit shapesTen class scores
Nobody wrote the edge detector — each layer is a feature stage the training data chose for itself.

The wall that hand-built features hit

The approach above has a name. It is called feature engineering, and for about thirty years it was the whole game in machine learning. The workflow looked like this: a human expert stared at the raw data, invented numerical summaries of it (the features), and fed those numbers to a relatively simple learning algorithm — a logistic regression, a support vector machine, a random forest.

That learning algorithm was genuinely learning. But it was only learning the last, easiest step. All the hard perceptual work — deciding what about the image matters — was done by the human in advance.

Here is what that looked like in code, using scikit-learn on our digits:

Python
import numpy as npfrom sklearn.linear_model import LogisticRegressiondef hand_built_features(image):    """Turn a 28x28 grey image into 5 numbers a human thought were relevant."""    ink = image > 0.5                      # boolean mask of "dark" pixels    rows, cols = np.nonzero(ink)    if len(rows) == 0:        return np.zeros(5)    return np.array([        ink.mean(),                        # how much ink overall        rows.mean() / 28.0,                # vertical centre of mass        cols.mean() / 28.0,                # horizontal centre of mass        (rows.max() - rows.min()) / 28.0,  # ink height        (cols.max() - cols.min()) / 28.0,  # ink width    ])X_features = np.array([hand_built_features(img) for img in train_images])clf = LogisticRegression(max_iter=1000).fit(X_features, train_labels)print(clf.score(np.array([hand_built_features(i) for i in test_images]), test_labels))# ~0.29 -- five human-chosen numbers simply do not contain enough information

Twenty-nine percent. Those five numbers throw away almost everything. You could spend a month inventing better ones and claw your way to the low nineties. That month is the bottleneck, and it does not transfer: features for handwriting tell you nothing about features for chest X-rays or for spoken Hindi.

Deep learning's central bet is that the feature-inventing step should itself be learned from data, not performed by a human in advance.

What "deep" actually means

A deep learning model is a stack of simple transformations, each one taking the output of the layer beneath it and producing a new representation for the layer above. "Deep" refers to the number of stacked layers. That is genuinely all the word means. It is not a claim about profundity.

Each layer does something almost embarrassingly simple: multiply the incoming numbers by a matrix of weights, add a vector of biases, then apply one non-linear function element by element. Written out, a three-layer network is:

h1=f(W1x+b1),h2=f(W2h1+b2),y^=W3h2+b3h_1 = f(W_1 x + b_1), \quad h_2 = f(W_2 h_1 + b_2), \quad \hat{y} = W_3 h_2 + b_3

Here xx is the input, the WW matrices and bb vectors are the numbers the model learns, ff is the non-linear function, and y^\hat{y} is the prediction. Stack enough of these and you get something that can read handwriting better than you can.

The same digits problem, done the deep way, needs no feature function at all:

Python
import torchimport torch.nn as nnmodel = nn.Sequential(    nn.Flatten(),            # 28x28 image -> 784 raw pixel values    nn.Linear(784, 256),     # learned transformation #1    nn.ReLU(),               # the non-linearity    nn.Linear(256, 128),     # learned transformation #2    nn.ReLU(),    nn.Linear(128, 10),      # one score per digit class)# Trained for ~5 epochs on 60,000 examples this reaches ~98% test accuracy.# Note what is absent: nobody told it about loops, strokes, or centres of mass.

Roughly 235,000 numbers get adjusted during training, and not one of them was designed by a person. The human contribution was the shape of the model, not its contents.

The non-linearity is not optional

Notice the nn.ReLU() between every pair of nn.Linear layers. Delete them and the model collapses. Two stacked matrix multiplications, W2(W1x)W_2(W_1x), are algebraically identical to a single matrix multiplication by W2W1W_2W_1. Without a non-linear function wedged between layers, a hundred-layer network has exactly the representational power of a one-layer network — which is to say, it can only draw straight lines through your data.

This is the single most important structural fact about neural networks, and it is why ReLU, which does nothing but max(0, x), is arguably the most consequential three lines of code in the field.

What the layers actually learn

If you train a network on images and then visualise what makes each unit fire hardest, a consistent pattern appears. It is one of the more beautiful empirical results in the subject.

DepthWhat that layer responds toAnalogy in handwriting
Layer 1Edges at particular orientations, blobs of light and dark"there is a stroke here, angled 30°"
Layer 2Corners, junctions, short curves — combinations of layer-1 edges"two strokes meet at a point"
Layer 3Recurring parts: loops, crossbars, closed regions"there is a closed loop in the upper half"
Layer 4+Whole-object configurations"loop on top, loop below, therefore 8"

The network rediscovered, on its own, the very features you were trying to hand-code — loops, crossbars, stroke junctions — and then found several hundred more that have no English name but work better than the ones that do.

This hierarchy explains why depth beats width. You could in principle solve the problem with one enormously wide layer; a result called the universal approximation theorem guarantees a single hidden layer with enough units can approximate any continuous function to arbitrary precision. But "enough units" can mean exponentially many. Depth lets the model reuse intermediate work: the same edge detector feeds a thousand different part detectors, which feed a hundred different object detectors. Composition is cheaper than enumeration.

A wide shallow network memorises combinations. A deep network builds vocabulary, then builds sentences out of it.

Why this took until 2012

The mathematics of multi-layer networks was worked out in the 1980s. Backpropagation, the algorithm that makes training possible, was published in 1986. And then almost nothing happened for twenty-five years. Understanding why is more useful than memorising the dates, because the same three constraints govern whether deep learning will work on your problem today.

IngredientSituation in 1990Situation nowWhy it mattered
Labelled dataA few thousand examples was a large datasetMillions of labelled images, billions of text tokensNetworks have millions of parameters; with few examples they memorise the training set instead of learning the pattern
ComputeCPUs doing a few million operations per secondGPUs doing tens of trillions of floating-point operations per secondTraining is billions of matrix multiplications; a run that takes 3 hours on a GPU takes weeks on a 1990 CPU
AlgorithmsSigmoid activations, naive random initialisationReLU, careful initialisation, batch normalisation, Adam, residual connectionsDeep stacks used to simply fail to train — gradients shrank to nothing before reaching the early layers

The 2012 moment was a network called AlexNet winning the ImageNet image-classification competition with a 15.3% error rate against the runner-up's 26.2%. That gap was not an incremental improvement; it was a different category of result, and it was produced by combining all three ingredients at once — 1.2 million labelled images, two consumer GPUs, and ReLU activations.

Where deep learning wins, and where it loses

This is where people waste the most money. Deep learning is not a general upgrade to machine learning. It is a specific tool with a specific profile, and on a large class of real business problems a gradient-boosted decision tree will beat a neural network while training in ninety seconds on a laptop.

SituationReach forReason
Images, audio, video, raw textDeep learningRaw signal with local structure; hand-built features are hopeless
Tabular data, < 100k rows, mixed typesGradient boosting (XGBoost, LightGBM)Consistently equal or better, far faster, far less tuning
Fewer than ~1,000 labelled examplesClassical models, or a pre-trained network fine-tunedTraining a deep model from scratch on tiny data overfits badly
The decision must be legally explainableLinear models, decision trees"The model's 40 million weights produced 0.83" is not an explanation a regulator accepts
Sequence-to-sequence: translation, transcription, generationDeep learningNo competitive alternative exists

Three costs deserve to be stated plainly, because they are routinely underestimated:

  • Data hunger. A rough working figure is a few thousand labelled examples per class before a from-scratch network becomes competitive. Labelling is usually the largest line item in the budget, not compute.
  • Opacity. When a network is confidently wrong, there is no line of code to inspect. Diagnosis means probing behaviour, not reading logic.
  • Brittleness to distribution shift. A model trained on photographs taken in daylight will quietly degrade on photographs taken at dusk, and it will not tell you it is degrading. It will keep returning confident probabilities.

The shape of every project

Whatever the domain, the loop is the same, and it is worth having it in your head before you write any code:

Text
1. Get data, split it     -> train / validation / test, split BEFORE any preprocessing2. Define an architecture -> how many layers, how wide, which non-linearity3. Define a loss          -> a single number measuring how wrong the model is4. Pick an optimiser      -> the rule for nudging weights to reduce that number5. Loop:     forward pass  -> compute predictions and loss     backward pass -> compute the gradient of loss w.r.t. every weight     update        -> step each weight a little way downhill6. Evaluate on validation -> diagnose underfitting vs overfitting7. Adjust and repeat      -> architecture, regularisation, learning rate8. Test once, at the end  -> the honest number you report

Step 8 carries a warning. The test set is spent the moment you look at it and change something in response. If you tune against your test set, the number you eventually publish is not an estimate of real-world performance — it is an estimate of how well you tuned. Keep it sealed.

What to do with this when you start building

The practical consequence of everything above is a triage habit. Before you open an editor, ask three questions in order.

Is the input raw perceptual signal? Pixels, waveforms, characters. If yes, deep learning is very likely the right tool, because the whole point is learning features from raw signal. If your input is forty tidy columns from a database, start with gradient boosting and treat a neural network as the thing you try when boosting has plateaued.

How many labelled examples exist today, not in the plan? If the honest answer is 300, no architecture will save you. Your options are to collect more, to fine-tune a model someone else trained on millions of examples, or to use a simpler method that survives on small data.

What happens when it is wrong? A misrouted letter gets re-sorted. A missed tumour does not. The tolerable error rate determines whether a 98%-accurate model is a triumph or a liability, and that decision belongs at the start of the project, not after the training run.

Get those three right and the rest of the work — choosing activations, writing training loops, chasing down a loss that has gone to NaN at 3am — is craft you can learn. Get them wrong and no amount of craft rescues the project.