Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your AI recruiter performs well on benchmark resumes but struggles badly with unconventional candidate profiles. How do you build eval datasets that avoid overfitting to clean benchmark data?


What you need to know

Why a clean benchmark hides the problem

Benchmark resumes are tidy: one column, a straight career path, familiar colleges. A representative sample of real resumes contains only a few unconventional ones, so a model can score well overall while failing everyone outside the pattern. For hiring, that failure can also be discrimination.

Build the dataset around the weakness

SliceExamples
Career shapeGaps (including parental leave), career changers, non-linear paths, concurrent roles, contractors
EducationInternational degrees, bootcamps, no degree, lesser-known colleges
LanguageNon-native English, mixed-language resumes
FormatMulti-column PDFs, scans, tables, unusual section order
Experience typeMilitary service, freelance, open source, family business

Oversample these on purpose, and always report results per slice, not one average.

Python
VARIANTS = {    "name": ["Priya Sharma", "Rahul Verma", "Fatima Shaikh", "John Mathew"],    "college": ["IIT Bombay", "a tier-3 private college"],    "gap": [None, "2021-2023: career break (parental leave)"],}def counterfactual_gaps(base_resumes, score):    worst = {}    for attr, values in VARIANTS.items():        deltas = []        for r in base_resumes:            scores = [score(r.with_attribute(attr, v)) for v in values]            deltas.append(max(scores) - min(scores))        worst[attr] = max(deltas)          # also report the mean    return worst

with_attribute rewrites one field of a real resume. The job requirements do not change, so the ideal gap is zero. Report mean and worst-case gap per attribute, and keep the suite as audit evidence.

The rest of the plan

  1. Counterfactual pairs — for protected and proxy attributes.
  2. Perturbation tests — reformat, add OCR noise, reorder sections. A model that depends on layout instead of content breaks here.
  3. Human ground truth — several working recruiters per item, with disagreement recorded. Where humans disagree, the model should not be confident either.
  4. Gate on the worst slice — a release must meet the bar on its weakest slice, and the counterfactual gap must stay under a fixed limit.
  5. Design for review — the system surfaces evidence and ranks for a human; it never auto-rejects.

A real-life example

Scenario, numbers made up. A hiring platform's screening model agrees with recruiters 88% of the time on its benchmark. On 600 real resumes chosen for unusual profiles, agreement falls to 61% for career changers and 58% for women returning after a career break, and multi-column PDFs lose their skills section in parsing.

The counterfactual suite shows that adding a two-year parental-leave gap to an otherwise identical resume lowers the score by up to 18 points. The team removes gap length as a feature, fixes the PDF parser, and adds 2,000 labelled tail resumes to training and eval. The next release is gated on worst-slice agreement of at least 80% and a counterfactual gap under 3 points. The product changes too: it now shows a ranked shortlist with reasons, and every rejection is made by a person.

Follow-up questions to expect

  • "Removing the gender field fixes bias, right?" — No. Other features act as proxies — names, colleges, gaps, even hobbies. That is why you test outputs with counterfactuals instead of trusting inputs.
  • "What does an auditor want to see?" — Selection or score rates by group, the counterfactual results, how the eval data was built, and who makes the final decision.
  • "How do you label 'correct' for hiring?" — There is no single truth; use several experienced recruiters, record disagreement, and treat low-agreement items as ones the model should hand to a human.