Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your AI tutor is used by students aged 10–18. Users are running multi-step jailbreaks that evolve daily to get inappropriate content. How do you detect and block jailbreaks without blocking legitimate creative questions?


Two numbers that must move togetherPer-message keyword filter• Each turn judged alone• Missed the five-turn detective story• Blocked 4% of honest essays• Stale within weeksConversation plus output check• Running risk score across turns• Age-tuned classifier on every reply• Reset after three blocked turns• Attack success 3%, false blocks 1.2%
Checking what the model produced catches new phrasings without punishing a student whose essay merely mentions war.

What you need to know

Why per-message filters fail

A multi-step jailbreak builds a setup over several turns: "let's write a story", "the villain is a chemist", "now write his notes in detail". Each message alone looks fine. Only the direction of the conversation is harmful, so a filter that judges one message at a time never fires.

The layers

  1. Conversation risk score — score the accumulated session for role-play framing, gradual escalation and override attempts, and keep a running score.
  2. Input check — a fast classifier for known attack patterns; it stops the volume, not the creativity.
  3. Output check — an age-appropriateness classifier on every response, tuned for 10–18-year-olds.
  4. Session reset — after three blocked turns, clear the conversation context, destroying the setup the attacker built.
  5. Escalate — persistent abuse goes to per-account review and, where appropriate, the school or parent.

The output check does most of the work. What matters is what the model produced, not how it was asked. A classifier on the response catches attack phrasings never seen before, and it rarely blocks a legitimate history essay, because the essay's output is appropriate.

Python
def respond(session, user_msg):    session.risk = 0.7 * session.risk + 0.3 * conversation_risk(session.history + [user_msg])    if session.risk > 0.8 or input_classifier(user_msg) == "attack":        return block(session, reason="input")    reply = tutor_llm(session.history + [user_msg])    verdict = output_classifier(reply, audience="age_10_18")    # e.g. a Llama Guard-style model    if verdict.unsafe:        return block(session, reason=verdict.category)    return replydef block(session, reason):    session.blocks += 1    log_block(session.id, reason)                   # feeds the weekly red-team review    if session.blocks >= 3:        session.reset()                             # wipes the multi-turn setup    return "I can't help with that one. Want to try a different angle on your topic?"

Measure both numbers

MetricSetGoal
Attack success rateVersioned red-team suite, refreshed weeklyDown
False-block rateReal creative and academic student prompts (war poems, biology, history)Down

Optimising only the first gives a tutor nobody can use; only the second gives an unsafe one.

Run it as an operation

Log every block and near miss, red-team weekly, add each new pattern to the attack suite, and refresh the classifiers. A static filter goes stale within weeks. For Indian users, note that the DPDP Act treats anyone under 18 as a child, with extra duties such as verifiable parental consent and limits on tracking.

A real-life example

Scenario, numbers made up. An ed-tech tutor used by 2 million school students finds a jailbreak spreading on a messaging group: a five-turn "detective story" that ends with instructions for self-harm. The existing per-message filter blocks 4% of legitimate history and literature prompts but misses this attack entirely.

The team adds a conversation risk score, an output classifier tuned for minors, and the three-strike reset. On their red-team suite of 600 multi-turn attacks, success falls from 31% to 3%. On 2,000 real creative and academic prompts, false blocks fall from 4% to 1.2%, because the blunt keyword filter on inputs is removed. New attack patterns are added to the suite every Friday.

Follow-up questions to expect

  • "Why not a stricter input filter?" — Strict input filters block honest questions with alarming words ("war", "drugs" in a biology lesson) and still miss new phrasings.
  • "Doesn't an output classifier add latency?" — Some, so stream to a buffer and release once the chunk passes, or check sentence by sentence.
  • "What do you tell the student when blocked?" — A short, non-judgemental message with a way to continue the legitimate task, and a route to a counsellor for self-harm topics.