AI Safety & Guardrails

Course Content

AI Safety & Guardrails

5 sections · 50 lessons

How do you prevent feedback loops that reinforce bias?


The screening loop that trains on its own choicesmodel shortlistsonlyshortlisted get hiredhiresbecome labelsretrain onthose labelsrejectednever get a labelTier-2 and tier-3 shortlist rate fell from 24% to 13% over four retrains.
The model never sees outcomes for people it rejected, so its mistakes become next quarter's ground truth.

What you need to know

Common loops

  • Lending: rejected applicants never get a chance to repay, so the model never learns it was wrong about them.
  • Hiring: only shortlisted candidates get interview scores or job performance data.
  • Recommendations: items shown more get more clicks, so they are shown more.
  • Fraud and policing: flagged groups are investigated more, so more confirmed cases come from them.
  • Model collapse: training on your own generated text narrows diversity over generations.

How to break them

  • Exploration: send a small random share (say 2–5%) through a path that ignores the model, or approve some borderline cases, so you observe outcomes the model would have suppressed. This has a real cost and needs ethical and business sign-off.
  • Off-policy correction: weight logged examples by the inverse of the probability that the old policy chose them, so training is less biased toward what the old model liked.
  • Segregate generated content: tag AI-generated text in your corpora so you can exclude or limit it.
  • Monitor per group over time: approval rates, score distributions and selection rates by cohort, across model versions.
  • Independent ground truth: a frozen, human-labelled benchmark the system has never influenced.
  • Review feedback before training: thumbs and clicks reflect what you showed, and can be gamed.

The fingerprint

A metric that drifts steadily in one direction across several model releases — for example, the selection rate for one group falling a little each retrain — is the signature of a feedback loop, even if each release looks fine alone.

A real-life example

A company's HR screening assistant is retrained every quarter on "successful hires", defined as employees rated "meets expectations" after a year. Only shortlisted candidates can become hires. After four retrains, the shortlist rate for candidates from tier-2 and tier-3 colleges falls from 24% to 13%, although their one-year ratings, when hired, match everyone else's.

The team: adds a 5% exploration slot where a recruiter reviews a random sample of rejected candidates, with those later hired entering the training data with the right weighting; freezes a 500-CV benchmark scored by a panel of hiring managers; and adds a release gate that blocks a new model if the shortlist rate for any college tier moves by more than 3 points without an explanation.

Follow-up questions to expect

  • "Isn't exploration unfair to the people in the random group?" — Design it so exploration only gives extra chances (reviewing rejected cases), not extra rejections. Get ethics and legal review.
  • "How is this different from data drift?" — Drift comes from the world changing; a feedback loop comes from your own system changing the data.
  • "Can you detect it without outcome labels for rejected cases?" — Partly: watch per-group selection trends across releases, and compare with an independent benchmark.