Course Content
AI Safety & Guardrails
5 sections · 50 lessons
How do you prevent misuse or abuse of AI systems in production?
What you need to know
"Misuse" covers two groups: abusers who use the system for harm (spam, scams, harassment, fraud) and attackers who target the system itself (extraction, cost attacks, jailbreaks). OWASP's LLM10: Unbounded Consumption covers the cost side: attacks that exhaust tokens, money or compute.
The layers
| Layer | What it stops |
|---|---|
| Identity | Anonymous bulk generation; gives you something to ban |
| Rate limits and quotas | Bulk abuse, extraction, scraping; per user, key, IP, org, with burst and sustained limits |
| Cost caps | Max output tokens, context size, timeouts, agent step and recursion limits |
| Moderation | Harmful content in and out, per policy category |
| Scope | Narrow prompt and few tools — less to misuse |
| Behavioural detection | Patterns across requests and accounts |
| Provenance | Labelling generated media (for example C2PA content credentials, watermarks) |
| Policy | An acceptable-use policy that you actually enforce |
Behavioural signals
Single requests often look innocent. Patterns do not:
- Sudden velocity spikes on a new account.
- High rates of "near misses" — requests just below a block threshold, or many blocked requests followed by a pass.
- Clusters of near-identical prompts across accounts.
- Several "different" users sharing a device, payment method or IP range.
Respond in steps: throttle, add a challenge, queue for review, suspend. Harsh first responses hurt legitimate users caught by mistake.
Measure it
- Abuse recall on a labelled sample of real traffic.
- False-positive rate for legitimate users.
- Cost per user distribution — abuse usually appears first as a long tail in spend.
A real-life example
A bank launches a free AI "loan eligibility explainer" on its website, no login needed. Within a week, scammers use it to generate hundreds of convincing fake loan-approval messages in the bank's style, and a competitor's script sends 50,000 questions a day to map the bank's eligibility rules.
The fixes: generation of letters or messages is removed from the public tool (scope); detailed questions need a logged-in session with OTP (identity); anonymous users get 10 questions per day per device (quota); output moderation blocks text formatted like official approval letters (content); and a dashboard tracks questions per session. Daily token spend falls 85%, and genuine usage, measured by logged-in customers who continue to an application, does not drop.
Follow-up questions to expect
- "What is a denial-of-wallet attack?" — Making your system spend money: long inputs, requests for maximum output, or loops that trigger many model calls. Token caps, budgets and step limits stop it.
- "How do you detect model extraction?" — High volumes of systematically varied queries from one account or a coordinated group, often covering the input space evenly instead of like real users.
- "How strict should limits be for paid enterprise customers?" — Higher and negotiated, but never unlimited; a compromised enterprise key is a common abuse source.