AI Safety & Guardrails

Course Content

AI Safety & Guardrails

5 sections · 50 lessons

How do you prevent misuse or abuse of AI systems in production?


What you need to know

"Misuse" covers two groups: abusers who use the system for harm (spam, scams, harassment, fraud) and attackers who target the system itself (extraction, cost attacks, jailbreaks). OWASP's LLM10: Unbounded Consumption covers the cost side: attacks that exhaust tokens, money or compute.

The layers

LayerWhat it stops
IdentityAnonymous bulk generation; gives you something to ban
Rate limits and quotasBulk abuse, extraction, scraping; per user, key, IP, org, with burst and sustained limits
Cost capsMax output tokens, context size, timeouts, agent step and recursion limits
ModerationHarmful content in and out, per policy category
ScopeNarrow prompt and few tools — less to misuse
Behavioural detectionPatterns across requests and accounts
ProvenanceLabelling generated media (for example C2PA content credentials, watermarks)
PolicyAn acceptable-use policy that you actually enforce

Behavioural signals

Single requests often look innocent. Patterns do not:

  • Sudden velocity spikes on a new account.
  • High rates of "near misses" — requests just below a block threshold, or many blocked requests followed by a pass.
  • Clusters of near-identical prompts across accounts.
  • Several "different" users sharing a device, payment method or IP range.

Respond in steps: throttle, add a challenge, queue for review, suspend. Harsh first responses hurt legitimate users caught by mistake.

Measure it

  • Abuse recall on a labelled sample of real traffic.
  • False-positive rate for legitimate users.
  • Cost per user distribution — abuse usually appears first as a long tail in spend.

A real-life example

A bank launches a free AI "loan eligibility explainer" on its website, no login needed. Within a week, scammers use it to generate hundreds of convincing fake loan-approval messages in the bank's style, and a competitor's script sends 50,000 questions a day to map the bank's eligibility rules.

The fixes: generation of letters or messages is removed from the public tool (scope); detailed questions need a logged-in session with OTP (identity); anonymous users get 10 questions per day per device (quota); output moderation blocks text formatted like official approval letters (content); and a dashboard tracks questions per session. Daily token spend falls 85%, and genuine usage, measured by logged-in customers who continue to an application, does not drop.

Follow-up questions to expect

  • "What is a denial-of-wallet attack?" — Making your system spend money: long inputs, requests for maximum output, or loops that trigger many model calls. Token caps, budgets and step limits stop it.
  • "How do you detect model extraction?" — High volumes of systematically varied queries from one account or a coordinated group, often covering the input space evenly instead of like real users.
  • "How strict should limits be for paid enterprise customers?" — Higher and negotiated, but never unlimited; a compromised enterprise key is a common abuse source.