Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your AI recruiter performs well on benchmark resumes but struggles badly with unconventional candidate profiles. How do you build eval datasets that avoid overfitting to clean benchmark data?
What you need to know
Why a clean benchmark hides the problem
Benchmark resumes are tidy: one column, a straight career path, familiar colleges. A representative sample of real resumes contains only a few unconventional ones, so a model can score well overall while failing everyone outside the pattern. For hiring, that failure can also be discrimination.
Build the dataset around the weakness
| Slice | Examples |
|---|---|
| Career shape | Gaps (including parental leave), career changers, non-linear paths, concurrent roles, contractors |
| Education | International degrees, bootcamps, no degree, lesser-known colleges |
| Language | Non-native English, mixed-language resumes |
| Format | Multi-column PDFs, scans, tables, unusual section order |
| Experience type | Military service, freelance, open source, family business |
Oversample these on purpose, and always report results per slice, not one average.
1VARIANTS = {2 "name": ["Priya Sharma", "Rahul Verma", "Fatima Shaikh", "John Mathew"],3 "college": ["IIT Bombay", "a tier-3 private college"],4 "gap": [None, "2021-2023: career break (parental leave)"],5}67def counterfactual_gaps(base_resumes, score):8 worst = {}9 for attr, values in VARIANTS.items():10 deltas = []11 for r in base_resumes:12 scores = [score(r.with_attribute(attr, v)) for v in values]13 deltas.append(max(scores) - min(scores))14 worst[attr] = max(deltas) # also report the mean15 return worstwith_attribute rewrites one field of a real resume. The job requirements do not change, so the ideal gap is zero. Report mean and worst-case gap per attribute, and keep the suite as audit evidence.
The rest of the plan
- Counterfactual pairs — for protected and proxy attributes.
- Perturbation tests — reformat, add OCR noise, reorder sections. A model that depends on layout instead of content breaks here.
- Human ground truth — several working recruiters per item, with disagreement recorded. Where humans disagree, the model should not be confident either.
- Gate on the worst slice — a release must meet the bar on its weakest slice, and the counterfactual gap must stay under a fixed limit.
- Design for review — the system surfaces evidence and ranks for a human; it never auto-rejects.
A real-life example
Scenario, numbers made up. A hiring platform's screening model agrees with recruiters 88% of the time on its benchmark. On 600 real resumes chosen for unusual profiles, agreement falls to 61% for career changers and 58% for women returning after a career break, and multi-column PDFs lose their skills section in parsing.
The counterfactual suite shows that adding a two-year parental-leave gap to an otherwise identical resume lowers the score by up to 18 points. The team removes gap length as a feature, fixes the PDF parser, and adds 2,000 labelled tail resumes to training and eval. The next release is gated on worst-slice agreement of at least 80% and a counterfactual gap under 3 points. The product changes too: it now shows a ranked shortlist with reasons, and every rejection is made by a person.
Follow-up questions to expect
- "Removing the gender field fixes bias, right?" — No. Other features act as proxies — names, colleges, gaps, even hobbies. That is why you test outputs with counterfactuals instead of trusting inputs.
- "What does an auditor want to see?" — Selection or score rates by group, the counterfactual results, how the eval data was built, and who makes the final decision.
- "How do you label 'correct' for hiring?" — There is no single truth; use several experienced recruiters, record disagreement, and treat low-agreement items as ones the model should hand to a human.