AI Safety & Guardrails

Course Content

AI Safety & Guardrails

5 sections · 50 lessons

How does differential privacy work in model training?


Per-example gradient norms in one batch0.220.3047.2012clippedto 1.0Clip every gradient to norm C = 1.0, sum, then add Gaussian noise scaled to C.
Clipping caps what any one record can do to the model; the noise then hides whether that record was there at all.

What you need to know

DP-SGD step by step

  1. Per-example gradients — compute the gradient for each training example separately, not just the batch average.
  2. Clip — scale any gradient whose length exceeds a limit C down to length C. This bounds how much one record can move the model.
  3. Add noise — add Gaussian noise with standard deviation sigma × C to the sum of clipped gradients.
  4. Average and update — divide by batch size and take the optimiser step.
  5. Account — a privacy accountant adds up the epsilon spent over every step; stop training when the budget is used.
Python
import math, randomdef clip(g, c):    norm = math.sqrt(sum(x * x for x in g))    return [x * min(1.0, c / norm) for x in g] if norm else gdef dp_step(per_example_grads, c=1.0, sigma=1.1):    clipped = [clip(g, c) for g in per_example_grads]    summed = [sum(col) for col in zip(*clipped)]    noisy = [s + random.gauss(0, sigma * c) for s in summed]    return [x / len(per_example_grads) for x in noisy]batch = [[0.2, -0.1], [0.3, 0.0], [25.0, 40.0]]   # last record is an outlierprint(dp_step(batch))

Without clipping, the outlier's gradient of length about 47 would dominate the update and could be detected. After clipping it counts no more than any other record, and the noise hides whether it was there at all. In practice you use a library: Opacus for PyTorch or TensorFlow Privacy, which handle per-example gradients and accounting.

Trade-offs

  • Accuracy drops, and it drops most for rare groups, because clipping and noise suppress exactly the unusual examples. DP can make fairness worse.
  • Compute rises: per-example gradients are slower and use more memory.
  • Tuning is sensitive: clipping norm, noise and batch size interact.
  • Epsilon values in published deployments range from below 1 to the low double digits; there is no universal "safe" number, so report it with delta and the unit of privacy (per example or per user).
  • For LLMs, DP fine-tuning with LoRA is the practical route, because fewer trainable parameters means less noise is needed.

What DP does not do

It protects the training set from being inferred from the model. It does nothing for data sent in a prompt at inference time, for data in a RAG index, or for logs. Those need redaction, access control and retention limits.

A real-life example

A hospital network wants a model that predicts readmission risk from 200,000 discharge records, and it plans to share the trained model with partner clinics. Without DP, a membership-inference test shows an attacker can tell with 64% accuracy (vs 50% for a coin flip) whether a specific patient was in training.

With DP-SGD at epsilon 3, the attack drops to 51%, and overall AUC falls from 0.81 to 0.78. But for patients over 80, a small group, AUC falls from 0.74 to 0.66. The team increases the batch size, adds more data from the older group, and reports per-group accuracy alongside epsilon in the model card.

Follow-up questions to expect

  • "What is a membership inference attack?" — Querying a model to decide whether a specific record was in its training data, often by checking whether the model is unusually confident on it. DP bounds how well this can work.
  • "Local vs central DP?" — Local DP adds noise on each user's device before data is collected; central DP trusts the collector and adds noise in the algorithm. Local is stronger but needs far more data for the same accuracy.
  • "Is synthetic data a replacement?" — Only if the synthetic-data generator was itself trained with DP. Otherwise it can copy real records.