Course Content
Deep Learning with TensorFlow and PyTorch
4 sections · 15 lessons
What Deep Learning Is and Why It Matters
Here is a job. A postal service hands you 60,000 scanned images of handwritten digits, each one a 28×28 grid of grey pixels, and asks you to write software that reads them. Get it wrong and letters go to the wrong city.
You are a competent programmer, so you start where competent programmers start: you look at the data and you think about features. A 7 has a long horizontal stroke at the top. An 8 has two enclosed loops. A 1 is mostly a single vertical line, so the ratio of ink-height to ink-width should be large. You write a function that counts enclosed regions. You write another that measures stroke thickness. You write a third that finds the centre of mass of the ink.
Then you look at the actual handwriting. Someone's 7 has a crossbar; someone else's does not. Someone's 4 is closed at the top and looks like a 9. Someone writes a 2 with a loop at the bottom, and now your enclosed-region counter says it is a 6. A 1 written at a slant has the same height-to-width ratio as a 7. You add exceptions. The exceptions collide with each other. After three weeks your accuracy has crawled to about 92%, which sounds respectable until you work out that it means roughly one letter in twelve gets misrouted.
The problem is not that you are bad at this. The problem is that you are trying to write down knowledge you do not consciously possess. You recognise a handwritten 7 in about 200 milliseconds and you have no idea how. The rules are not in your head as rules. They are in your head as something else, and the entire field of deep learning exists because someone worked out how to build that something else out of arithmetic.
The wall that hand-built features hit
The approach above has a name. It is called feature engineering, and for about thirty years it was the whole game in machine learning. The workflow looked like this: a human expert stared at the raw data, invented numerical summaries of it (the features), and fed those numbers to a relatively simple learning algorithm — a logistic regression, a support vector machine, a random forest.
That learning algorithm was genuinely learning. But it was only learning the last, easiest step. All the hard perceptual work — deciding what about the image matters — was done by the human in advance.
Here is what that looked like in code, using scikit-learn on our digits:
1import numpy as np2from sklearn.linear_model import LogisticRegression34def hand_built_features(image):5 """Turn a 28x28 grey image into 5 numbers a human thought were relevant."""6 ink = image > 0.5 # boolean mask of "dark" pixels7 rows, cols = np.nonzero(ink)8 if len(rows) == 0:9 return np.zeros(5)10 return np.array([11 ink.mean(), # how much ink overall12 rows.mean() / 28.0, # vertical centre of mass13 cols.mean() / 28.0, # horizontal centre of mass14 (rows.max() - rows.min()) / 28.0, # ink height15 (cols.max() - cols.min()) / 28.0, # ink width16 ])1718X_features = np.array([hand_built_features(img) for img in train_images])19clf = LogisticRegression(max_iter=1000).fit(X_features, train_labels)20print(clf.score(np.array([hand_built_features(i) for i in test_images]), test_labels))21# ~0.29 -- five human-chosen numbers simply do not contain enough informationTwenty-nine percent. Those five numbers throw away almost everything. You could spend a month inventing better ones and claw your way to the low nineties. That month is the bottleneck, and it does not transfer: features for handwriting tell you nothing about features for chest X-rays or for spoken Hindi.
Deep learning's central bet is that the feature-inventing step should itself be learned from data, not performed by a human in advance.
What "deep" actually means
A deep learning model is a stack of simple transformations, each one taking the output of the layer beneath it and producing a new representation for the layer above. "Deep" refers to the number of stacked layers. That is genuinely all the word means. It is not a claim about profundity.
Each layer does something almost embarrassingly simple: multiply the incoming numbers by a matrix of weights, add a vector of biases, then apply one non-linear function element by element. Written out, a three-layer network is:
Here x is the input, the W matrices and b vectors are the numbers the model learns, f is the non-linear function, and y^ is the prediction. Stack enough of these and you get something that can read handwriting better than you can.
The same digits problem, done the deep way, needs no feature function at all:
1import torch2import torch.nn as nn34model = nn.Sequential(5 nn.Flatten(), # 28x28 image -> 784 raw pixel values6 nn.Linear(784, 256), # learned transformation #17 nn.ReLU(), # the non-linearity8 nn.Linear(256, 128), # learned transformation #29 nn.ReLU(),10 nn.Linear(128, 10), # one score per digit class11)12# Trained for ~5 epochs on 60,000 examples this reaches ~98% test accuracy.13# Note what is absent: nobody told it about loops, strokes, or centres of mass.Roughly 235,000 numbers get adjusted during training, and not one of them was designed by a person. The human contribution was the shape of the model, not its contents.
The non-linearity is not optional
Notice the nn.ReLU() between every pair of nn.Linear layers. Delete them and the model collapses. Two stacked matrix multiplications, W2(W1x), are algebraically identical to a single matrix multiplication by W2W1. Without a non-linear function wedged between layers, a hundred-layer network has exactly the representational power of a one-layer network — which is to say, it can only draw straight lines through your data.
This is the single most important structural fact about neural networks, and it is why ReLU, which does nothing but max(0, x), is arguably the most consequential three lines of code in the field.
What the layers actually learn
If you train a network on images and then visualise what makes each unit fire hardest, a consistent pattern appears. It is one of the more beautiful empirical results in the subject.
| Depth | What that layer responds to | Analogy in handwriting |
|---|---|---|
| Layer 1 | Edges at particular orientations, blobs of light and dark | "there is a stroke here, angled 30°" |
| Layer 2 | Corners, junctions, short curves — combinations of layer-1 edges | "two strokes meet at a point" |
| Layer 3 | Recurring parts: loops, crossbars, closed regions | "there is a closed loop in the upper half" |
| Layer 4+ | Whole-object configurations | "loop on top, loop below, therefore 8" |
The network rediscovered, on its own, the very features you were trying to hand-code — loops, crossbars, stroke junctions — and then found several hundred more that have no English name but work better than the ones that do.
This hierarchy explains why depth beats width. You could in principle solve the problem with one enormously wide layer; a result called the universal approximation theorem guarantees a single hidden layer with enough units can approximate any continuous function to arbitrary precision. But "enough units" can mean exponentially many. Depth lets the model reuse intermediate work: the same edge detector feeds a thousand different part detectors, which feed a hundred different object detectors. Composition is cheaper than enumeration.
A wide shallow network memorises combinations. A deep network builds vocabulary, then builds sentences out of it.
Why this took until 2012
The mathematics of multi-layer networks was worked out in the 1980s. Backpropagation, the algorithm that makes training possible, was published in 1986. And then almost nothing happened for twenty-five years. Understanding why is more useful than memorising the dates, because the same three constraints govern whether deep learning will work on your problem today.
| Ingredient | Situation in 1990 | Situation now | Why it mattered |
|---|---|---|---|
| Labelled data | A few thousand examples was a large dataset | Millions of labelled images, billions of text tokens | Networks have millions of parameters; with few examples they memorise the training set instead of learning the pattern |
| Compute | CPUs doing a few million operations per second | GPUs doing tens of trillions of floating-point operations per second | Training is billions of matrix multiplications; a run that takes 3 hours on a GPU takes weeks on a 1990 CPU |
| Algorithms | Sigmoid activations, naive random initialisation | ReLU, careful initialisation, batch normalisation, Adam, residual connections | Deep stacks used to simply fail to train — gradients shrank to nothing before reaching the early layers |
The 2012 moment was a network called AlexNet winning the ImageNet image-classification competition with a 15.3% error rate against the runner-up's 26.2%. That gap was not an incremental improvement; it was a different category of result, and it was produced by combining all three ingredients at once — 1.2 million labelled images, two consumer GPUs, and ReLU activations.
Where deep learning wins, and where it loses
This is where people waste the most money. Deep learning is not a general upgrade to machine learning. It is a specific tool with a specific profile, and on a large class of real business problems a gradient-boosted decision tree will beat a neural network while training in ninety seconds on a laptop.
| Situation | Reach for | Reason |
|---|---|---|
| Images, audio, video, raw text | Deep learning | Raw signal with local structure; hand-built features are hopeless |
| Tabular data, < 100k rows, mixed types | Gradient boosting (XGBoost, LightGBM) | Consistently equal or better, far faster, far less tuning |
| Fewer than ~1,000 labelled examples | Classical models, or a pre-trained network fine-tuned | Training a deep model from scratch on tiny data overfits badly |
| The decision must be legally explainable | Linear models, decision trees | "The model's 40 million weights produced 0.83" is not an explanation a regulator accepts |
| Sequence-to-sequence: translation, transcription, generation | Deep learning | No competitive alternative exists |
Three costs deserve to be stated plainly, because they are routinely underestimated:
- Data hunger. A rough working figure is a few thousand labelled examples per class before a from-scratch network becomes competitive. Labelling is usually the largest line item in the budget, not compute.
- Opacity. When a network is confidently wrong, there is no line of code to inspect. Diagnosis means probing behaviour, not reading logic.
- Brittleness to distribution shift. A model trained on photographs taken in daylight will quietly degrade on photographs taken at dusk, and it will not tell you it is degrading. It will keep returning confident probabilities.
The shape of every project
Whatever the domain, the loop is the same, and it is worth having it in your head before you write any code:
1. Get data, split it -> train / validation / test, split BEFORE any preprocessing2. Define an architecture -> how many layers, how wide, which non-linearity3. Define a loss -> a single number measuring how wrong the model is4. Pick an optimiser -> the rule for nudging weights to reduce that number5. Loop: forward pass -> compute predictions and loss backward pass -> compute the gradient of loss w.r.t. every weight update -> step each weight a little way downhill6. Evaluate on validation -> diagnose underfitting vs overfitting7. Adjust and repeat -> architecture, regularisation, learning rate8. Test once, at the end -> the honest number you reportStep 8 carries a warning. The test set is spent the moment you look at it and change something in response. If you tune against your test set, the number you eventually publish is not an estimate of real-world performance — it is an estimate of how well you tuned. Keep it sealed.
What to do with this when you start building
The practical consequence of everything above is a triage habit. Before you open an editor, ask three questions in order.
Is the input raw perceptual signal? Pixels, waveforms, characters. If yes, deep learning is very likely the right tool, because the whole point is learning features from raw signal. If your input is forty tidy columns from a database, start with gradient boosting and treat a neural network as the thing you try when boosting has plateaued.
How many labelled examples exist today, not in the plan? If the honest answer is 300, no architecture will save you. Your options are to collect more, to fine-tune a model someone else trained on millions of examples, or to use a simpler method that survives on small data.
What happens when it is wrong? A misrouted letter gets re-sorted. A missed tumour does not. The tolerable error rate determines whether a 98%-accurate model is a triumph or a liability, and that decision belongs at the start of the project, not after the training run.
Get those three right and the rest of the work — choosing activations, writing training loops, chasing down a loss that has gone to NaN at 3am — is craft you can learn. Get them wrong and no amount of craft rescues the project.