Deep Learning with TensorFlow and PyTorch

Setting Up TensorFlow and PyTorch


Almost everyone's first day with a deep learning framework goes the same way. You type pip install tensorflow, wait through a few hundred megabytes of download, open Python, and get this:

Text
>>> import tensorflow as tf>>> tf.config.list_physical_devices('GPU')[]

An empty list. You have a GPU. You paid for the GPU. TensorFlow cannot see it. Somewhere in the install log, buried under warnings, is a line about libcudart.so.12 not being found.

Then you fix that, and two days later you pip install something unrelated — a plotting library, a data tool — and it quietly upgrades NumPy, and now import torch raises an error about a binary incompatibility in a module you have never heard of.

Neither of these is a deep learning problem. Both are dependency-management problems, and they eat more beginner time than backpropagation ever does. The habits below cost twenty minutes to establish and save you days.

The install that actually works on day oneOne fresh envper projectInstall the buildmatching your CUDAVerify: versionand device countMove modeland data toone deviceSeed everything,then record itA tensor on CPU and a model on GPU is the first error nearly everyone meets.
The import succeeding proves nothing; only a tensor that lands on the GPU proves the install.

Why one environment per project is not optional

These frameworks pin their dependencies tightly, because they ship compiled C++ and CUDA code that is built against specific versions of NumPy, protobuf and the CUDA runtime. Install two projects into the same Python and you eventually get an unsatisfiable constraint:

Text
Project A needs:  tensorflow  -> numpy<2.1Project B needs:  some-tool   -> numpy>=2.2pip installs B, silently upgrades numpy, and now:  import tensorflow  ImportError: numpy.core.multiarray failed to import

The fix is one environment per project. Any of the three standard tools will do it:

Bash
# Option 1: built into Python, nothing to installpython -m venv .venvsource .venv/bin/activate          # Windows: .venv\Scripts\activate# Option 2: conda, which also manages non-Python libraries like CUDAconda create -n dl python=3.11 -yconda activate dl# Option 3: uv, considerably faster at resolving and installinguv venv --python 3.11source .venv/bin/activate

Whichever you choose, the discipline is the same: activate before installing, and record what you installed.

Bash
pip freeze > requirements.txt      # exact versions, for reproducing this run later

An experiment you cannot reproduce is an anecdote. The version file is part of the experiment, not paperwork about it.

Installing PyTorch

PyTorch ships separate wheels for each accelerator backend, and the download URL selects which one you get. This is the part people get wrong: plain pip install torch may give you a build with no GPU support at all, depending on your platform.

Bash
# NVIDIA GPU on Linux or Windows -- pick the CUDA build matching your driver# (pytorch.org/get-started lists the builds offered for the current release)pip install torch torchvision --index-url https://download.pytorch.org/whl/cu130# CPU only -- much smaller download, fine for learning and small modelspip install torch torchvision --index-url https://download.pytorch.org/whl/cpu# Apple Silicon (M-series) -- the default wheel includes the Metal backendpip install torch torchvision

Check your NVIDIA driver first with nvidia-smi. The "CUDA Version" it reports is the highest CUDA runtime your driver supports; the PyTorch wheel's CUDA version must be at or below it. A driver reporting 13.0 runs a cu126 wheel happily; the reverse fails.

Installing TensorFlow

Bash
# Linux with NVIDIA GPU: the extra pulls in the CUDA libraries as pip packagespip install "tensorflow[and-cuda]"# CPU only, any platformpip install tensorflow# Apple Silicon: base package plus the Metal pluginpip install tensorflow tensorflow-metal

On Apple Silicon, check the tensorflow-metal release notes before you upgrade TensorFlow: the plugin is released separately and has lagged behind recent TensorFlow versions, so a new TensorFlow can leave you with no GPU until the plugin catches up.

The [and-cuda] extra matters. Without it TensorFlow expects a system-wide CUDA installation with the right versions on your library path, which is exactly the configuration that produces the empty-GPU-list problem. With it, the CUDA runtime arrives as ordinary Python packages inside your environment, which is both easier and properly isolated.

One platform note that saves a lot of confusion: TensorFlow has not shipped native GPU support on Windows since version 2.10. On Windows you either use WSL2 (Windows Subsystem for Linux, where the Linux instructions apply unchanged) or accept CPU-only training.

Verifying the install properly

"It imported" is not verification. Run something that touches the accelerator and returns a number.

Python
import torchprint("torch:", torch.__version__)print("CUDA available:", torch.cuda.is_available())if torch.cuda.is_available():    print("device:", torch.cuda.get_device_name(0))    free, total = torch.cuda.mem_get_info()    print(f"memory: {free/1e9:.1f} GB free of {total/1e9:.1f} GB")print("MPS available:", torch.backends.mps.is_available())   # Apple Silicon# Actually compute something on the devicedevice = ("cuda" if torch.cuda.is_available()          else "mps" if torch.backends.mps.is_available() else "cpu")a = torch.randn(2000, 2000, device=device)print("matmul ok:", (a @ a).sum().item())
Python
import tensorflow as tfprint("tensorflow:", tf.__version__)print("GPUs:", tf.config.list_physical_devices("GPU"))with tf.device("/GPU:0" if tf.config.list_physical_devices("GPU") else "/CPU:0"):    a = tf.random.normal((2000, 2000))    print("matmul ok:", float(tf.reduce_sum(a @ a)))

If the matrix multiply runs and returns a finite number, the install is real.

One TensorFlow setting worth applying immediately

By default TensorFlow grabs essentially all GPU memory at start-up, which means a second process (a Jupyter kernel you forgot about, a colleague on a shared machine) will fail with an out-of-memory error even when there is plenty free. Turn on incremental growth before you touch anything else:

Python
import tensorflow as tffor gpu in tf.config.list_physical_devices("GPU"):    tf.config.experimental.set_memory_growth(gpu, True)# Must run BEFORE any tensor is created or any model is built.

Tensors: the shared vocabulary

Both frameworks are built on the same object — an n-dimensional array of numbers that knows which device it lives on and can be differentiated through. Only the spelling differs.

OperationPyTorchTensorFlow
From a Python listtorch.tensor([1., 2.])tf.constant([1., 2.])
Zeros of a given shapetorch.zeros(3, 4)tf.zeros((3, 4))
Random normaltorch.randn(3, 4)tf.random.normal((3, 4))
Shapex.shapex.shape
Reshapex.view(2, 6) or x.reshape(2, 6)tf.reshape(x, (2, 6))
Matrix multiplya @ ba @ b
Add an axisx.unsqueeze(0)tf.expand_dims(x, 0)
To NumPyx.detach().cpu().numpy()x.numpy()
Trainable variabletorch.tensor(..., requires_grad=True)tf.Variable(...)

The x.detach().cpu().numpy() incantation in PyTorch is worth unpacking, since you will type it constantly. detach() removes the tensor from the autograd graph so NumPy conversion does not error; cpu() moves it off the GPU, because NumPy cannot read GPU memory; numpy() does the conversion. Skip any step and you get an error message that names exactly the step you skipped.

Devices, and the error you will hit today

A GPU has its own memory, separate from system RAM. An operation needs all its inputs in the same place. This produces the single most common runtime error in PyTorch:

Text
RuntimeError: Expected all tensors to be on the same device,but found at least two devices, cuda:0 and cpu!

The cure is a discipline: choose the device once, move the model once, and move every batch as it comes out of the loader.

Python
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")model = MyModel().to(device)              # move parameters once, at constructionfor xb, yb in train_loader:    xb, yb = xb.to(device), yb.to(device) # move each batch, every iteration    loss = loss_fn(model(xb), yb)

Two traps hide in that snippet. First, model.to(device) modifies the module in place and also returns it, but tensor.to(device) does not modify in place — it returns a new tensor. Writing xb.to(device) without assigning the result is a no-op that produces exactly the error above. Second, build your optimiser after moving the model. In current PyTorch model.to(device) keeps the same parameter objects, so the other order usually works too, but this is the order the documentation recommends and some backends (TPUs through PyTorch/XLA, for example) do need it.

TensorFlow places tensors on the GPU automatically when one is visible, so this class of error is rarer there. The trade-off is less control: when you need it, use an explicit with tf.device("/CPU:0"): block.

The same model, written twice

Here is a small classifier in both frameworks, doing exactly the same arithmetic. Reading them side by side is the fastest way to internalise the difference in philosophy.

Python
import tensorflow as tffrom tensorflow import kerasmodel = keras.Sequential([    keras.layers.Input(shape=(20,)),    keras.layers.Dense(64, activation="relu"),    keras.layers.Dense(32, activation="relu"),    keras.layers.Dense(1),])model.compile(    optimizer=keras.optimizers.Adam(1e-3),    loss=keras.losses.BinaryCrossentropy(from_logits=True),    metrics=["accuracy"],)model.fit(X_train, y_train, epochs=20, batch_size=32,          validation_data=(X_val, y_val))
Python
import torchimport torch.nn as nnmodel = nn.Sequential(    nn.Linear(20, 64), nn.ReLU(),    nn.Linear(64, 32), nn.ReLU(),    nn.Linear(32, 1),).to(device)opt = torch.optim.Adam(model.parameters(), lr=1e-3)loss_fn = nn.BCEWithLogitsLoss()for epoch in range(20):    model.train()    for xb, yb in train_loader:        xb, yb = xb.to(device), yb.to(device)        opt.zero_grad()                       # clear accumulated gradients        loss = loss_fn(model(xb), yb)         # forward        loss.backward()                       # backward        opt.step()                            # update    model.eval()    with torch.no_grad():        val_loss = sum(loss_fn(model(xb.to(device)), yb.to(device)).item()                       for xb, yb in val_loader) / len(val_loader)    print(f"epoch {epoch}: val_loss={val_loss:.4f}")

Keras gives you fit(), which is twelve lines of well-tested machinery you did not have to write. PyTorch makes you write those twelve lines, which means you can see and change every one of them. Neither is better in general; they suit different moments.

What actually differs

DimensionTensorFlow / KerasPyTorch
Default executionEager, with @tf.function to compile a graphEager, with torch.compile() to compile
Training loopfit() provided; GradientTape when you need customYou write it; libraries like Lightning wrap it
DebuggingStraightforward in eager, harder inside @tf.functionPlain Python; pdb and print work anywhere
DeploymentMature: TF Serving, LiteRT (formerly TF Lite), TensorFlow.jstorch.export, ONNX export, ExecuTorch; general servers such as NVIDIA Triton (TorchServe is archived)
Research code you will find onlineLess common nowDominant — most papers ship PyTorch
Mobile and embeddedLiteRT (formerly TF Lite) is very well establishedExecuTorch, younger but now production-ready

A common misconception worth dispelling: "TensorFlow is static graphs, PyTorch is dynamic". That was true in 2017. TensorFlow 2 runs eagerly by default and PyTorch has a graph compiler. Today the meaningful differences are ergonomic and ecosystem-related, not fundamental.

Making runs reproducible

Weight initialisation, data shuffling, and dropout masks are all random. Two runs of identical code produce different numbers, and without seeding you cannot tell whether your architecture change helped or you got lucky.

Python
import os, random, numpy as np, torchdef set_seed(seed=42):    random.seed(seed)    np.random.seed(seed)    torch.manual_seed(seed)    torch.cuda.manual_seed_all(seed)    os.environ["PYTHONHASHSEED"] = str(seed)    # Fully deterministic GPU kernels -- slower, but bit-identical runs    torch.backends.cudnn.deterministic = True    torch.backends.cudnn.benchmark = Falseset_seed(42)
Python
import tensorflow as tftf.keras.utils.set_random_seed(42)   # seeds Python, NumPy and TF togethertf.config.experimental.enable_op_determinism()

Be honest about what this buys you. Determinism costs perhaps 10–20% throughput, and it does not survive a change of GPU model or library version. Its real value is comparing two of your own runs on the same machine on the same day — which is exactly what you do all day when tuning.

The errors you will meet in week one

MessageCauseFix
list_physical_devices('GPU') returns []CPU-only wheel, or missing CUDA librariesReinstall with tensorflow[and-cuda]; check nvidia-smi works
Expected all tensors to be on the same deviceModel on GPU, batch still on CPU (or vice versa)xb = xb.to(device) — assign the result
CUDA out of memoryBatch too large, or an unreleased graph being heldSmaller batch; wrap evaluation in torch.no_grad(); accumulate loss.item() not loss
mat1 and mat2 shapes cannot be multipliedLayer input size does not match the incoming tensorPrint x.shape between layers; check for a missing Flatten
numpy.core.multiarray failed to importNumPy upgraded out from under a compiled packageRebuild the environment; pin versions in requirements.txt
Loss is nan from step oneLearning rate far too high, or unnormalised inputsDrop the LR by 10×; standardise features to roughly zero mean, unit variance
Training runs but nothing improvesOptimiser holds different parameters from the model (built for an earlier copy, or layers kept in a plain list), or opt.step() never calledBuild the optimiser from the model you train; print a parameter before and after a step to confirm it moves

A setup checklist worth keeping

Before you start any new project, run through this. It takes five minutes and prevents most of the above.

  1. Create a fresh environment and activate it. Never install into the system Python.
  2. Install the framework with the accelerator variant that matches your hardware, not the default wheel.
  3. Run a real computation on the device — a matrix multiply that returns a finite number — not just an import.
  4. Enable GPU memory growth if you are using TensorFlow, before creating anything.
  5. Set seeds at the top of every script.
  6. Freeze versions to requirements.txt and commit it alongside the code.
  7. Train one tiny model end to end — twenty examples, three epochs — before scaling up. If the pipeline is broken, find out in ten seconds rather than after an hour of GPU time.

That last point is the one people skip and regret. The purpose of the first training run is not to get a good model; it is to prove that data flows in, gradients flow back, and a checkpoint file appears on disk. Once that loop is verified, everything after it is tuning.