Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Attackers discover that adding invisible Unicode characters bypasses your moderation filters. How do you design robust moderation systems against adversarial prompt attacks?


What you need to know

Why invisible characters work

A zero-width space (U+200B) shows nothing on screen. Put one inside a banned word and a keyword filter no longer matches, while the language model still reads the word easily. Other tricks work the same way: Cyrillic "а" in place of Latin "a" (they look identical), full-width letters, letters separated by spaces, or text wrapped in base64 with "decode this".

Normalise a copy, then classify

Python
import re, unicodedataINVISIBLE = re.compile(r"[\u200b\u200e\u200f\u202a-\u202e\u2060-\u2064\ufeff\u00ad\U000e0000-\U000e007f]")CONFUSABLES = {"а": "a", "е": "e", "о": "o", "р": "p", "с": "c", "і": "i"}  # sample; load the full UTS 39 datadef canonical(text: str) -> tuple[str, int]:    t = unicodedata.normalize("NFKC", text)        # full-width and styled letters -> plain    hidden = len(INVISIBLE.findall(t))    t = INVISIBLE.sub("", t)    t = "".join(CONFUSABLES.get(ch, ch) for ch in t)    return t, hiddenclean, hidden = canonical(user_text)risk = classifier.score(clean) + (0.3 if hidden else 0.0)

NFKC turns full-width and "mathematical bold" letters into plain ones. It does not turn Cyrillic into Latin, which is why a confusables table (Unicode Technical Standard 39) is needed. The range U+E0000–E007F covers "tag" characters, which are invisible and have been used to smuggle hidden instructions to models. The zero-width joiner and non-joiner (U+200C, U+200D) are left out on purpose: see the next section.

The normalised copy is only for checking. Store and display the user's original text.

Don't break real languages

Zero-width joiners are legitimate in emoji sequences and in several Indic scripts, such as Malayalam and Devanagari. Stripping them everywhere, or treating them as abuse, flags normal Hindi, Malayalam and emoji messages. Handle them by context — suspicious inside a Latin-script word, normal inside Indic text or emoji — and measure false positives on real non-English traffic before enforcing.

The layers that matter most

  1. Canonicalise at the edge — normalise, strip, map, decode wrappers.
  2. Score the obfuscation — ordinary users do not hide zero-width characters inside English words.
  3. Use a trained classifier — a small fine-tuned model generalises across obfuscation families; regex lists are always one trick behind.
  4. Moderate the output — the model's reply is plain text, and it is what causes harm. Output checks survive input encodings you have never seen.
  5. Run a mutation harness — apply obfuscation transforms to a known-bad set every day and track the bypass rate as a metric.

A real-life example

Scenario, numbers made up. A social app's AI chat feature uses a keyword list plus a toxicity model on raw input. A post on a forum shows that inserting zero-width spaces gets abusive requests through. The mutation harness, built that week, confirms a 34% bypass rate across 5,000 known-bad prompts.

The team adds canonicalisation, a hidden-character risk feature and output moderation. Bypass falls to 3%. The first rollout, which also stripped zero-width joiners, flagged 2% of Malayalam messages as suspicious; after switching to context-aware handling, false positives on Indic-language traffic return to the old baseline.

Follow-up questions to expect

  • "Why not just block any message with invisible characters?" — Legitimate text uses some of them (emoji, Indic scripts). Treat them as a risk signal, not an automatic block.
  • "What about new tricks you haven't seen?" — Output moderation catches harm regardless of how the input was encoded, and the mutation harness grows as new tricks appear.
  • "Where does this run?" — At the gateway, before any model call, so every product using the model gets the same protection.