Course Content
AI Safety & Guardrails
5 sections · 50 lessons
How do you handle proxy variables that indirectly encode sensitive attributes?
What you need to know
Common proxies
| Proxy | Can encode |
|---|---|
| PIN code / neighbourhood | Caste, religion, income, ethnicity |
| First name, surname | Gender, religion, community |
| College, school | Gender (women's colleges), region, income |
| Career gaps | Parenthood, illness, disability |
| Language, dialect, writing style | Region, origin, education |
| Device type, payment method | Income |
In a text model, the whole input is a proxy. A CV's writing style, hobbies and clubs can reveal gender and background.
Measuring proxy strength
1from collections import Counter, defaultdict23def proxy_strength(rows, feature, sensitive):4 by_value = defaultdict(Counter)5 for r in rows:6 by_value[r[feature]][r[sensitive]] += 17 hits = sum(c.most_common(1)[0][1] for c in by_value.values())8 baseline = Counter(r[sensitive] for r in rows).most_common(1)[0][1]9 return round(hits / len(rows), 2), round(baseline / len(rows), 2)1011rows = ([{"college": "A", "gender": "F"}] * 90 + [{"college": "A", "gender": "M"}] * 10 +12 [{"college": "B", "gender": "M"}] * 160 + [{"college": "B", "gender": "F"}] * 40)13print(proxy_strength(rows, "college", "gender")) # (0.83, 0.57)Guessing gender from college alone is right 83% of the time, against 57% for always guessing the majority. College is a strong proxy. On real data you train a model on all features to predict the sensitive attribute and look at its AUC and the most important features.
What to do about it
- Coarsen or drop proxies that add little real predictive value — and measure the accuracy cost, and for whom.
- Counterfactual tests: change only the proxy (swap the name, the PIN code) and measure how much the output moves.
- Training constraints: adversarial debiasing, fairness-constrained optimisation, reweighting.
- Judge by outcomes: selection rates, error-rate gaps and calibration by group on real decisions.
- Measure with the attribute: keep it separate from training features, access-controlled, used only for fairness measurement.
A real-life example
A lender's credit model does not use religion or caste. An audit groups applicants by PIN code and finds that approval rates in 40 PIN codes with mostly minority populations are 18 points lower than similar-income areas. The PIN-code feature, plus "distance to branch", carries most of that difference.
The team replaces six-digit PIN code with a district-level cost-of-living index, removes distance to branch (which added almost no predictive value once income was known), and re-runs the audit. The gap falls to 5 points; overall default-prediction AUC drops from 0.79 to 0.78. They document the trade-off, and monitor the gap every month.
Follow-up questions to expect
- "Why not just remove every correlated feature?" — Almost every useful feature correlates with something sensitive. Remove those that add little real signal; for the rest, constrain and measure.
- "Is collecting sensitive attributes legal?" — Often yes for fairness monitoring with a proper basis, consent and strict access control, but check local law; in some places it is restricted, which is why this needs legal review.
- "How do proxies show up in LLMs?" — The model infers gender or background from names, writing style or details and changes its tone or recommendation. Counterfactual name-swap tests catch this.