Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your AI tutor is used by students aged 10–18. Users are running multi-step jailbreaks that evolve daily to get inappropriate content. How do you detect and block jailbreaks without blocking legitimate creative questions?
What you need to know
Why per-message filters fail
A multi-step jailbreak builds a setup over several turns: "let's write a story", "the villain is a chemist", "now write his notes in detail". Each message alone looks fine. Only the direction of the conversation is harmful, so a filter that judges one message at a time never fires.
The layers
- Conversation risk score — score the accumulated session for role-play framing, gradual escalation and override attempts, and keep a running score.
- Input check — a fast classifier for known attack patterns; it stops the volume, not the creativity.
- Output check — an age-appropriateness classifier on every response, tuned for 10–18-year-olds.
- Session reset — after three blocked turns, clear the conversation context, destroying the setup the attacker built.
- Escalate — persistent abuse goes to per-account review and, where appropriate, the school or parent.
The output check does most of the work. What matters is what the model produced, not how it was asked. A classifier on the response catches attack phrasings never seen before, and it rarely blocks a legitimate history essay, because the essay's output is appropriate.
1def respond(session, user_msg):2 session.risk = 0.7 * session.risk + 0.3 * conversation_risk(session.history + [user_msg])3 if session.risk > 0.8 or input_classifier(user_msg) == "attack":4 return block(session, reason="input")5 reply = tutor_llm(session.history + [user_msg])6 verdict = output_classifier(reply, audience="age_10_18") # e.g. a Llama Guard-style model7 if verdict.unsafe:8 return block(session, reason=verdict.category)9 return reply1011def block(session, reason):12 session.blocks += 113 log_block(session.id, reason) # feeds the weekly red-team review14 if session.blocks >= 3:15 session.reset() # wipes the multi-turn setup16 return "I can't help with that one. Want to try a different angle on your topic?"Measure both numbers
| Metric | Set | Goal |
|---|---|---|
| Attack success rate | Versioned red-team suite, refreshed weekly | Down |
| False-block rate | Real creative and academic student prompts (war poems, biology, history) | Down |
Optimising only the first gives a tutor nobody can use; only the second gives an unsafe one.
Run it as an operation
Log every block and near miss, red-team weekly, add each new pattern to the attack suite, and refresh the classifiers. A static filter goes stale within weeks. For Indian users, note that the DPDP Act treats anyone under 18 as a child, with extra duties such as verifiable parental consent and limits on tracking.
A real-life example
Scenario, numbers made up. An ed-tech tutor used by 2 million school students finds a jailbreak spreading on a messaging group: a five-turn "detective story" that ends with instructions for self-harm. The existing per-message filter blocks 4% of legitimate history and literature prompts but misses this attack entirely.
The team adds a conversation risk score, an output classifier tuned for minors, and the three-strike reset. On their red-team suite of 600 multi-turn attacks, success falls from 31% to 3%. On 2,000 real creative and academic prompts, false blocks fall from 4% to 1.2%, because the blunt keyword filter on inputs is removed. New attack patterns are added to the suite every Friday.
Follow-up questions to expect
- "Why not a stricter input filter?" — Strict input filters block honest questions with alarming words ("war", "drugs" in a biology lesson) and still miss new phrasings.
- "Doesn't an output classifier add latency?" — Some, so stream to a buffer and release once the chunk passes, or check sentence by sentence.
- "What do you tell the student when blocked?" — A short, non-judgemental message with a way to continue the legitimate task, and a route to a counsellor for self-harm topics.