Course Content
Deep Learning with TensorFlow and PyTorch
4 sections · 15 lessons
Setting Up TensorFlow and PyTorch
Almost everyone's first day with a deep learning framework goes the same way. You type pip install tensorflow, wait through a few hundred megabytes of download, open Python, and get this:
>>> import tensorflow as tf>>> tf.config.list_physical_devices('GPU')[]An empty list. You have a GPU. You paid for the GPU. TensorFlow cannot see it. Somewhere in the install log, buried under warnings, is a line about libcudart.so.12 not being found.
Then you fix that, and two days later you pip install something unrelated — a plotting library, a data tool — and it quietly upgrades NumPy, and now import torch raises an error about a binary incompatibility in a module you have never heard of.
Neither of these is a deep learning problem. Both are dependency-management problems, and they eat more beginner time than backpropagation ever does. The habits below cost twenty minutes to establish and save you days.
Why one environment per project is not optional
These frameworks pin their dependencies tightly, because they ship compiled C++ and CUDA code that is built against specific versions of NumPy, protobuf and the CUDA runtime. Install two projects into the same Python and you eventually get an unsatisfiable constraint:
Project A needs: tensorflow -> numpy<2.1Project B needs: some-tool -> numpy>=2.2pip installs B, silently upgrades numpy, and now: import tensorflow ImportError: numpy.core.multiarray failed to importThe fix is one environment per project. Any of the three standard tools will do it:
1# Option 1: built into Python, nothing to install2python -m venv .venv3source .venv/bin/activate # Windows: .venv\Scripts\activate45# Option 2: conda, which also manages non-Python libraries like CUDA6conda create -n dl python=3.11 -y7conda activate dl89# Option 3: uv, considerably faster at resolving and installing10uv venv --python 3.1111source .venv/bin/activateWhichever you choose, the discipline is the same: activate before installing, and record what you installed.
pip freeze > requirements.txt # exact versions, for reproducing this run laterAn experiment you cannot reproduce is an anecdote. The version file is part of the experiment, not paperwork about it.
Installing PyTorch
PyTorch ships separate wheels for each accelerator backend, and the download URL selects which one you get. This is the part people get wrong: plain pip install torch may give you a build with no GPU support at all, depending on your platform.
1# NVIDIA GPU on Linux or Windows -- pick the CUDA build matching your driver2# (pytorch.org/get-started lists the builds offered for the current release)3pip install torch torchvision --index-url https://download.pytorch.org/whl/cu13045# CPU only -- much smaller download, fine for learning and small models6pip install torch torchvision --index-url https://download.pytorch.org/whl/cpu78# Apple Silicon (M-series) -- the default wheel includes the Metal backend9pip install torch torchvisionCheck your NVIDIA driver first with nvidia-smi. The "CUDA Version" it reports is the highest CUDA runtime your driver supports; the PyTorch wheel's CUDA version must be at or below it. A driver reporting 13.0 runs a cu126 wheel happily; the reverse fails.
Installing TensorFlow
1# Linux with NVIDIA GPU: the extra pulls in the CUDA libraries as pip packages2pip install "tensorflow[and-cuda]"34# CPU only, any platform5pip install tensorflow67# Apple Silicon: base package plus the Metal plugin8pip install tensorflow tensorflow-metalOn Apple Silicon, check the tensorflow-metal release notes before you upgrade TensorFlow: the plugin is released separately and has lagged behind recent TensorFlow versions, so a new TensorFlow can leave you with no GPU until the plugin catches up.
The [and-cuda] extra matters. Without it TensorFlow expects a system-wide CUDA installation with the right versions on your library path, which is exactly the configuration that produces the empty-GPU-list problem. With it, the CUDA runtime arrives as ordinary Python packages inside your environment, which is both easier and properly isolated.
One platform note that saves a lot of confusion: TensorFlow has not shipped native GPU support on Windows since version 2.10. On Windows you either use WSL2 (Windows Subsystem for Linux, where the Linux instructions apply unchanged) or accept CPU-only training.
Verifying the install properly
"It imported" is not verification. Run something that touches the accelerator and returns a number.
1import torch23print("torch:", torch.__version__)4print("CUDA available:", torch.cuda.is_available())5if torch.cuda.is_available():6 print("device:", torch.cuda.get_device_name(0))7 free, total = torch.cuda.mem_get_info()8 print(f"memory: {free/1e9:.1f} GB free of {total/1e9:.1f} GB")9print("MPS available:", torch.backends.mps.is_available()) # Apple Silicon1011# Actually compute something on the device12device = ("cuda" if torch.cuda.is_available()13 else "mps" if torch.backends.mps.is_available() else "cpu")14a = torch.randn(2000, 2000, device=device)15print("matmul ok:", (a @ a).sum().item())1import tensorflow as tf23print("tensorflow:", tf.__version__)4print("GPUs:", tf.config.list_physical_devices("GPU"))56with tf.device("/GPU:0" if tf.config.list_physical_devices("GPU") else "/CPU:0"):7 a = tf.random.normal((2000, 2000))8 print("matmul ok:", float(tf.reduce_sum(a @ a)))If the matrix multiply runs and returns a finite number, the install is real.
One TensorFlow setting worth applying immediately
By default TensorFlow grabs essentially all GPU memory at start-up, which means a second process (a Jupyter kernel you forgot about, a colleague on a shared machine) will fail with an out-of-memory error even when there is plenty free. Turn on incremental growth before you touch anything else:
1import tensorflow as tf23for gpu in tf.config.list_physical_devices("GPU"):4 tf.config.experimental.set_memory_growth(gpu, True)5# Must run BEFORE any tensor is created or any model is built.Tensors: the shared vocabulary
Both frameworks are built on the same object — an n-dimensional array of numbers that knows which device it lives on and can be differentiated through. Only the spelling differs.
| Operation | PyTorch | TensorFlow |
|---|---|---|
| From a Python list | torch.tensor([1., 2.]) | tf.constant([1., 2.]) |
| Zeros of a given shape | torch.zeros(3, 4) | tf.zeros((3, 4)) |
| Random normal | torch.randn(3, 4) | tf.random.normal((3, 4)) |
| Shape | x.shape | x.shape |
| Reshape | x.view(2, 6) or x.reshape(2, 6) | tf.reshape(x, (2, 6)) |
| Matrix multiply | a @ b | a @ b |
| Add an axis | x.unsqueeze(0) | tf.expand_dims(x, 0) |
| To NumPy | x.detach().cpu().numpy() | x.numpy() |
| Trainable variable | torch.tensor(..., requires_grad=True) | tf.Variable(...) |
The x.detach().cpu().numpy() incantation in PyTorch is worth unpacking, since you will type it constantly. detach() removes the tensor from the autograd graph so NumPy conversion does not error; cpu() moves it off the GPU, because NumPy cannot read GPU memory; numpy() does the conversion. Skip any step and you get an error message that names exactly the step you skipped.
Devices, and the error you will hit today
A GPU has its own memory, separate from system RAM. An operation needs all its inputs in the same place. This produces the single most common runtime error in PyTorch:
RuntimeError: Expected all tensors to be on the same device,but found at least two devices, cuda:0 and cpu!The cure is a discipline: choose the device once, move the model once, and move every batch as it comes out of the loader.
1device = torch.device("cuda" if torch.cuda.is_available() else "cpu")2model = MyModel().to(device) # move parameters once, at construction34for xb, yb in train_loader:5 xb, yb = xb.to(device), yb.to(device) # move each batch, every iteration6 loss = loss_fn(model(xb), yb)Two traps hide in that snippet. First, model.to(device) modifies the module in place and also returns it, but tensor.to(device) does not modify in place — it returns a new tensor. Writing xb.to(device) without assigning the result is a no-op that produces exactly the error above. Second, build your optimiser after moving the model. In current PyTorch model.to(device) keeps the same parameter objects, so the other order usually works too, but this is the order the documentation recommends and some backends (TPUs through PyTorch/XLA, for example) do need it.
TensorFlow places tensors on the GPU automatically when one is visible, so this class of error is rarer there. The trade-off is less control: when you need it, use an explicit with tf.device("/CPU:0"): block.
The same model, written twice
Here is a small classifier in both frameworks, doing exactly the same arithmetic. Reading them side by side is the fastest way to internalise the difference in philosophy.
1import tensorflow as tf2from tensorflow import keras34model = keras.Sequential([5 keras.layers.Input(shape=(20,)),6 keras.layers.Dense(64, activation="relu"),7 keras.layers.Dense(32, activation="relu"),8 keras.layers.Dense(1),9])1011model.compile(12 optimizer=keras.optimizers.Adam(1e-3),13 loss=keras.losses.BinaryCrossentropy(from_logits=True),14 metrics=["accuracy"],15)16model.fit(X_train, y_train, epochs=20, batch_size=32,17 validation_data=(X_val, y_val))1import torch2import torch.nn as nn34model = nn.Sequential(5 nn.Linear(20, 64), nn.ReLU(),6 nn.Linear(64, 32), nn.ReLU(),7 nn.Linear(32, 1),8).to(device)910opt = torch.optim.Adam(model.parameters(), lr=1e-3)11loss_fn = nn.BCEWithLogitsLoss()1213for epoch in range(20):14 model.train()15 for xb, yb in train_loader:16 xb, yb = xb.to(device), yb.to(device)17 opt.zero_grad() # clear accumulated gradients18 loss = loss_fn(model(xb), yb) # forward19 loss.backward() # backward20 opt.step() # update2122 model.eval()23 with torch.no_grad():24 val_loss = sum(loss_fn(model(xb.to(device)), yb.to(device)).item()25 for xb, yb in val_loader) / len(val_loader)26 print(f"epoch {epoch}: val_loss={val_loss:.4f}")Keras gives you fit(), which is twelve lines of well-tested machinery you did not have to write. PyTorch makes you write those twelve lines, which means you can see and change every one of them. Neither is better in general; they suit different moments.
What actually differs
| Dimension | TensorFlow / Keras | PyTorch |
|---|---|---|
| Default execution | Eager, with @tf.function to compile a graph | Eager, with torch.compile() to compile |
| Training loop | fit() provided; GradientTape when you need custom | You write it; libraries like Lightning wrap it |
| Debugging | Straightforward in eager, harder inside @tf.function | Plain Python; pdb and print work anywhere |
| Deployment | Mature: TF Serving, LiteRT (formerly TF Lite), TensorFlow.js | torch.export, ONNX export, ExecuTorch; general servers such as NVIDIA Triton (TorchServe is archived) |
| Research code you will find online | Less common now | Dominant — most papers ship PyTorch |
| Mobile and embedded | LiteRT (formerly TF Lite) is very well established | ExecuTorch, younger but now production-ready |
A common misconception worth dispelling: "TensorFlow is static graphs, PyTorch is dynamic". That was true in 2017. TensorFlow 2 runs eagerly by default and PyTorch has a graph compiler. Today the meaningful differences are ergonomic and ecosystem-related, not fundamental.
Making runs reproducible
Weight initialisation, data shuffling, and dropout masks are all random. Two runs of identical code produce different numbers, and without seeding you cannot tell whether your architecture change helped or you got lucky.
1import os, random, numpy as np, torch23def set_seed(seed=42):4 random.seed(seed)5 np.random.seed(seed)6 torch.manual_seed(seed)7 torch.cuda.manual_seed_all(seed)8 os.environ["PYTHONHASHSEED"] = str(seed)9 # Fully deterministic GPU kernels -- slower, but bit-identical runs10 torch.backends.cudnn.deterministic = True11 torch.backends.cudnn.benchmark = False1213set_seed(42)import tensorflow as tftf.keras.utils.set_random_seed(42) # seeds Python, NumPy and TF togethertf.config.experimental.enable_op_determinism()Be honest about what this buys you. Determinism costs perhaps 10–20% throughput, and it does not survive a change of GPU model or library version. Its real value is comparing two of your own runs on the same machine on the same day — which is exactly what you do all day when tuning.
The errors you will meet in week one
| Message | Cause | Fix |
|---|---|---|
list_physical_devices('GPU') returns [] | CPU-only wheel, or missing CUDA libraries | Reinstall with tensorflow[and-cuda]; check nvidia-smi works |
Expected all tensors to be on the same device | Model on GPU, batch still on CPU (or vice versa) | xb = xb.to(device) — assign the result |
CUDA out of memory | Batch too large, or an unreleased graph being held | Smaller batch; wrap evaluation in torch.no_grad(); accumulate loss.item() not loss |
mat1 and mat2 shapes cannot be multiplied | Layer input size does not match the incoming tensor | Print x.shape between layers; check for a missing Flatten |
numpy.core.multiarray failed to import | NumPy upgraded out from under a compiled package | Rebuild the environment; pin versions in requirements.txt |
Loss is nan from step one | Learning rate far too high, or unnormalised inputs | Drop the LR by 10×; standardise features to roughly zero mean, unit variance |
| Training runs but nothing improves | Optimiser holds different parameters from the model (built for an earlier copy, or layers kept in a plain list), or opt.step() never called | Build the optimiser from the model you train; print a parameter before and after a step to confirm it moves |
A setup checklist worth keeping
Before you start any new project, run through this. It takes five minutes and prevents most of the above.
- Create a fresh environment and activate it. Never install into the system Python.
- Install the framework with the accelerator variant that matches your hardware, not the default wheel.
- Run a real computation on the device — a matrix multiply that returns a finite number — not just an import.
- Enable GPU memory growth if you are using TensorFlow, before creating anything.
- Set seeds at the top of every script.
- Freeze versions to
requirements.txtand commit it alongside the code. - Train one tiny model end to end — twenty examples, three epochs — before scaling up. If the pipeline is broken, find out in ten seconds rather than after an hour of GPU time.
That last point is the one people skip and regret. The purpose of the first training run is not to get a good model; it is to prove that data flows in, gradients flow back, and a checkpoint file appears on disk. Once that loop is verified, everything after it is tuning.