Course Content
AI Safety & Guardrails
5 sections · 50 lessons
How do you prevent feedback loops that reinforce bias?
What you need to know
Common loops
- Lending: rejected applicants never get a chance to repay, so the model never learns it was wrong about them.
- Hiring: only shortlisted candidates get interview scores or job performance data.
- Recommendations: items shown more get more clicks, so they are shown more.
- Fraud and policing: flagged groups are investigated more, so more confirmed cases come from them.
- Model collapse: training on your own generated text narrows diversity over generations.
How to break them
- Exploration: send a small random share (say 2–5%) through a path that ignores the model, or approve some borderline cases, so you observe outcomes the model would have suppressed. This has a real cost and needs ethical and business sign-off.
- Off-policy correction: weight logged examples by the inverse of the probability that the old policy chose them, so training is less biased toward what the old model liked.
- Segregate generated content: tag AI-generated text in your corpora so you can exclude or limit it.
- Monitor per group over time: approval rates, score distributions and selection rates by cohort, across model versions.
- Independent ground truth: a frozen, human-labelled benchmark the system has never influenced.
- Review feedback before training: thumbs and clicks reflect what you showed, and can be gamed.
The fingerprint
A metric that drifts steadily in one direction across several model releases — for example, the selection rate for one group falling a little each retrain — is the signature of a feedback loop, even if each release looks fine alone.
A real-life example
A company's HR screening assistant is retrained every quarter on "successful hires", defined as employees rated "meets expectations" after a year. Only shortlisted candidates can become hires. After four retrains, the shortlist rate for candidates from tier-2 and tier-3 colleges falls from 24% to 13%, although their one-year ratings, when hired, match everyone else's.
The team: adds a 5% exploration slot where a recruiter reviews a random sample of rejected candidates, with those later hired entering the training data with the right weighting; freezes a 500-CV benchmark scored by a panel of hiring managers; and adds a release gate that blocks a new model if the shortlist rate for any college tier moves by more than 3 points without an explanation.
Follow-up questions to expect
- "Isn't exploration unfair to the people in the random group?" — Design it so exploration only gives extra chances (reviewing rejected cases), not extra rejections. Get ethics and legal review.
- "How is this different from data drift?" — Drift comes from the world changing; a feedback loop comes from your own system changing the data.
- "Can you detect it without outcome labels for rejected cases?" — Partly: watch per-group selection trends across releases, and compare with an independent benchmark.