Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your public chatbot API is being abused by bots generating spam and phishing content at scale. How do you implement abuse detection and rate limiting for LLM APIs?


What you need to know

What is being stolen

The abuser wants your model's output: thousands of phishing emails, fake reviews or spam posts, paid for by you. So measure and limit output tokens and spend, not only request count.

Layered limits

LayerExample
Per keyRequests per minute and output tokens per minute
Per accountDaily spend cap; lower caps for new accounts on probation
Per payment method or phoneCaps shared across all accounts that use it
GlobalA circuit breaker if total generation spikes
Python
import time, redisr = redis.Redis()def allow_output(account: str, tokens: int, per_minute: int, per_day: int) -> bool:    minute = f"otok:{account}:{int(time.time() // 60)}"    day = f"otok:{account}:{time.strftime('%Y%m%d')}"    pipe = r.pipeline()    pipe.incrby(minute, tokens); pipe.expire(minute, 120)    pipe.incrby(day, tokens);    pipe.expire(day, 172800)    used_min, _, used_day, _ = pipe.execute()    return used_min <= per_minute and used_day <= per_day

These are fixed windows counted in Redis: simple and enough to start. A sliding window is smoother at the edges. Charge the tokens after generation, and use the running total to decide whether the next request is allowed.

Behaviour gives bots away

Spam traffic has a shape:

  • Near-duplicate prompts at volume, often a template with swapped names or brands (MinHash finds these).
  • Flat traffic all day and night, with no human daily rhythm.
  • High output-to-input ratio, single-turn only, no edits or follow-ups.
  • Many new accounts sharing a payment method, device fingerprint, IP range or prompt template.

Cluster accounts on these features to catch farms that spread across many keys.

Classify the actual harm

A general toxicity classifier passes a polite, well-written phishing email. Train or use classifiers for the specific harms: phishing and credential harvesting, bulk marketing spam, brand impersonation, fake reviews.

  1. Make identity cost something — verified sign-up, email-domain reputation, card or phone above the free tier, probation limits for new accounts.
  2. Limit tokens and spend — per key, account and payment method.
  3. Detect behaviour — duplicate clustering and traffic-shape features.
  4. Classify harms — phishing, spam, impersonation.
  5. Respond gradually — quietly slow down first, then challenge (CAPTCHA or re-verification), then suspend. An instant hard block shows attackers exactly where your limit is.

Track abuse prevalence in sampled traffic, time to detect a new campaign, and false positives on paying customers — the most expensive error.

A real-life example

Scenario, numbers made up. A startup's public chatbot API offers a free tier. In one week, output tokens rise 9 times. A sample shows fake "KYC update" emails pretending to be from Indian banks, each asking the reader to click a link.

Clustering finds 2,300 free accounts sending prompts from four templates, all created in 48 hours, with no activity gaps day or night. The team suspends the cluster, adds phone verification above 20 requests a day, sets daily output-token caps for accounts younger than a week, and adds a phishing classifier on outputs. Free-tier cost falls back to normal within two days. Two paying customers with bulk marketing use cases trip the duplicate detector in the first week and are moved to an allowlist after review.

Follow-up questions to expect

  • "Won't shadow-throttling hurt real users?" — It is applied only to accounts with high risk scores, and slowing down is easy to reverse. Watch the false-positive rate on paying customers closely.
  • "Why not just block IPs?" — Abusers rotate through residential proxies and cloud IPs; IP is one weak signal among many, not the control.
  • "How do you catch a new campaign fast?" — Alert on sudden new clusters of near-duplicate prompts and on output-token growth per cohort of new accounts.