AI Safety & Guardrails

Course Content

AI Safety & Guardrails

5 sections · 50 lessons

How do you adapt moderation systems across different cultures?


Moderation policy, from fixed to localUniversal floor:CSAM, terror, threatsCore policy,thresholds by localeLocal policy withregional expertsJurisdictionlayer owned by legaltopbottomHinglish recall rose from 41% to 83% with native-speaker eval sets.
Separating the fixed floor from local layers lets policy adapt without one country's law becoming everyone's rule.

What you need to know

Why one global system fails

  • Context: a word that is a slur in one region is neutral in another; caste-based abuse is a major harm category in South Asia that global taxonomies often under-cover.
  • Language quality: classifiers trained mostly on English perform much worse in low-resource languages.
  • Code-mixing and transliteration: Hinglish ("yeh banda bilkul bekaar hai"), Tamil written in Latin script, Arabic written in numbers-and-letters — all confuse monolingual models.
  • Law: countries differ on what must be removed and how fast. India's IT Rules, for example, set grievance and takedown obligations for intermediaries.

A layered policy

LayerContentConfigurable?
Universal floorChild sexual abuse material, terrorist content, credible threatsNo
Core policyHate, harassment, self-harm, violenceThresholds by locale
Local policyRegion-specific slurs, caste abuse, local sensitivitiesYes, defined with regional experts
Jurisdiction layerLegally mandated removals per countryLegal-owned, separate

Keeping the jurisdiction layer separate means one country's legal requirement does not quietly become global policy.

Evaluation

  • Native-speaker labels from the region, not translations of an English set. Translation loses exactly the context that matters.
  • Per-language and per-script metrics: aggregate multilingual numbers hide failures.
  • False-positive rate per community: classifiers often over-flag minority dialects and reclaimed in-group language. That over-enforcement is itself a fairness harm.
  • Launch gate: "we have not measured recall in this language" blocks launch there.

Tools

Many provider moderation APIs and open models like Llama Guard support several languages, but coverage and accuracy vary; check the model card's language list and test on your own data.

A real-life example

A food-delivery app adds AI moderation to restaurant reviews in India. The English classifier catches 88% of abusive English reviews but only 41% of abusive Hinglish reviews, and it flags 12% of harmless Tamil-script reviews because it has barely seen Tamil.

The team builds evaluation sets of 2,000 reviews each for English, Hinglish, Hindi (Devanagari), Tamil and Bengali, labelled by native-speaker moderators; adds a transliteration step and a multilingual classifier; adds a caste-slur list maintained with regional experts; and routes low-confidence non-English cases to regional human reviewers. Recall for Hinglish reaches 83%, and false flags on Tamil fall to 2%.

Follow-up questions to expect

  • "Should you translate everything to English and moderate that?" — It is a quick baseline, but translation drops slang, tone and cultural meaning, so accuracy suffers most on the cases that matter.
  • "How do you handle reclaimed words?" — Use context and community-aware labels; measure false positives on in-group speech specifically.
  • "Who writes the local policy?" — Regional policy experts and native-speaker reviewers, versioned and reviewed like code, with a central team keeping consistency.