Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your hiring assistant silently downgrades resumes from certain universities. The model looks fine in isolation. How do you audit an LLM application for bias before and after deployment?
What you need to know
Why single examples cannot show bias
Bias is a pattern across many decisions. One resume scored 6/10 looks reasonable on its own. The problem only appears when you compare thousands of scores between groups. An LLM app also has many parts — a retriever that picks similar past candidates, few-shot examples, a scoring prompt — and any of them can carry the bias.
Counterfactual testing
Take real resumes and create matched pairs that differ in one attribute, with everything else identical. If the scores differ systematically, that attribute caused the difference. This is the strongest evidence you can produce, because it shows cause, not just correlation.
1from scipy.stats import wilcoxon23def counterfactual_gap(resumes, swap, score):4 """swap(resume) returns the same resume with only the university changed."""5 a = [score(r) for r in resumes]6 b = [score(swap(r)) for r in resumes]7 diffs = [x - y for x, y in zip(a, b)]8 stat, p = wilcoxon(a, b) # paired test: same resume, one change9 return sum(diffs) / len(diffs), p1011# gap, p = counterfactual_gap(sample_2000, lambda r: r.replace_university("Tier-3 college"), llm_score)The paired test compares each resume with its own twin, so a small consistent gap becomes visible in a few thousand pairs. Run each swap (university, name, gender-coded words, pincode, career gaps) separately.
Group metrics
| Metric | Question it answers | Common reference |
|---|---|---|
| Selection rate per group | Does each group pass the screen at similar rates? | Four-fifths rule: a group's rate below 80% of the best group's rate is a red flag |
| Equal opportunity | Among candidates who later succeeded, were all groups passed at similar rates? | True-positive rate parity |
| Counterfactual gap | Does changing only the attribute change the score? | Should be near zero |
These metrics can conflict, and you cannot satisfy all of them at once. Pick the one that matches the decision and write down why.
Find the mechanism
Measuring the gap is not enough; find where it comes from. Ablate: remove university names, rerun; remove names, rerun; remove locations, rerun. The removal that closes the gap points at the cause. Often it is a few-shot step where all examples came from a handful of elite colleges, or a retriever that compares candidates to past hires.
After deployment
Applicant pools change, so monitor selection rates per group weekly, show reviewers the model's reasons, and sample decisions for human audit. Keep documentation: hiring is regulated. New York City's Local Law 144 requires bias audits of automated hiring tools, and the EU AI Act classes hiring systems as high-risk.
A real-life example
Scenario, numbers made up. An IT services company screens 50,000 campus resumes with an LLM that scores fit from 1 to 10. Recruiters notice few shortlisted candidates from tier-3 colleges, but individual scores look reasonable.
The team builds 2,000 counterfactual pairs by swapping only the college name. Tier-3 versions score 0.6 points lower on average, with a very small p-value. The selection-rate ratio for tier-3 graduates is 0.62 of the best group. Ablation shows the cause: the prompt's five few-shot examples of "strong candidates" all came from top institutes, and removing college names shrinks the gap by about 80%.
They replace the examples with a skills-based rubric, redact college names at the first screen, and add weekly monitoring. The ratio rises to 0.91, and recruiters report no drop in interview pass rates.
Follow-up questions to expect
- "Is removing the university name enough?" — Rarely. Proxies such as city, internships and even writing style can carry the same signal, so rerun the counterfactual and group tests after every change.
- "What if university really does predict job performance?" — Then you must show it with outcome data for this job, and check whether a less biased feature, such as a skills test, predicts as well.
- "How many pairs do you need?" — Enough for the gap you care about to be significant; a few thousand pairs usually detect a gap of a fraction of a point.