Course Content
AI Safety & Guardrails
5 sections · 50 lessons
How do you adapt moderation systems across different cultures?
What you need to know
Why one global system fails
- Context: a word that is a slur in one region is neutral in another; caste-based abuse is a major harm category in South Asia that global taxonomies often under-cover.
- Language quality: classifiers trained mostly on English perform much worse in low-resource languages.
- Code-mixing and transliteration: Hinglish ("yeh banda bilkul bekaar hai"), Tamil written in Latin script, Arabic written in numbers-and-letters — all confuse monolingual models.
- Law: countries differ on what must be removed and how fast. India's IT Rules, for example, set grievance and takedown obligations for intermediaries.
A layered policy
| Layer | Content | Configurable? |
|---|---|---|
| Universal floor | Child sexual abuse material, terrorist content, credible threats | No |
| Core policy | Hate, harassment, self-harm, violence | Thresholds by locale |
| Local policy | Region-specific slurs, caste abuse, local sensitivities | Yes, defined with regional experts |
| Jurisdiction layer | Legally mandated removals per country | Legal-owned, separate |
Keeping the jurisdiction layer separate means one country's legal requirement does not quietly become global policy.
Evaluation
- Native-speaker labels from the region, not translations of an English set. Translation loses exactly the context that matters.
- Per-language and per-script metrics: aggregate multilingual numbers hide failures.
- False-positive rate per community: classifiers often over-flag minority dialects and reclaimed in-group language. That over-enforcement is itself a fairness harm.
- Launch gate: "we have not measured recall in this language" blocks launch there.
Tools
Many provider moderation APIs and open models like Llama Guard support several languages, but coverage and accuracy vary; check the model card's language list and test on your own data.
A real-life example
A food-delivery app adds AI moderation to restaurant reviews in India. The English classifier catches 88% of abusive English reviews but only 41% of abusive Hinglish reviews, and it flags 12% of harmless Tamil-script reviews because it has barely seen Tamil.
The team builds evaluation sets of 2,000 reviews each for English, Hinglish, Hindi (Devanagari), Tamil and Bengali, labelled by native-speaker moderators; adds a transliteration step and a multilingual classifier; adds a caste-slur list maintained with regional experts; and routes low-confidence non-English cases to regional human reviewers. Recall for Hinglish reaches 83%, and false flags on Tamil fall to 2%.
Follow-up questions to expect
- "Should you translate everything to English and moderate that?" — It is a quick baseline, but translation drops slang, tone and cultural meaning, so accuracy suffers most on the cases that matter.
- "How do you handle reclaimed words?" — Use context and community-aware labels; measure false positives on in-group speech specifically.
- "Who writes the local policy?" — Regional policy experts and native-speaker reviewers, versioned and reviewed like code, with a central team keeping consistency.