Course Content
AI Safety & Guardrails
5 sections · 50 lessons
What software engineering practices improve AI system reliability?
What you need to know
Version everything
- Prompts, tool schemas and eval sets in git, reviewed like code.
- Pinned model IDs with a date or version, never a "latest" alias that can change under you.
- Index snapshots with a version, so you know which documents a given answer used.
Test like you mean it
- Eval gate in CI: a golden dataset with thresholds; the build fails if quality drops.
- Deterministic unit tests: record model responses as fixtures so tests of your code don't depend on the model.
- Regression cases: every production failure becomes a test.
Handle bad outputs as normal events
- Structured outputs with a schema; validate, and retry once with the error message if it fails.
- Treat a malformed output as a caught error, not an exception that crashes a downstream job.
Resilience patterns
| Pattern | Why |
|---|---|
| Timeouts | Model calls can hang for tens of seconds |
| Retries with backoff and jitter | Handle rate limits and transient errors |
| Circuit breaker | Stop hammering a failing provider |
| Fallback | Smaller model, cached answer, or "try later" |
| Idempotency keys | A retried issue_refund must not refund twice |
Release safely
- Feature flags per capability.
- Shadow traffic: run the new version on real requests without showing users.
- Canary: 1–5% of traffic, compare metrics, then expand.
- One-click rollback.
Observe
- Tracing of prompt, retrieved context, tool calls, model version, tokens, latency and cost per request (OpenTelemetry has semantic conventions for generative AI; tools like Langfuse and LangSmith build on the same idea).
- Quality alerts: refusal rate, guardrail block rate, groundedness score, cost per request, p95 latency — not only 5xx errors.
A real-life example
A food-delivery company's support assistant is quietly switched to a newer model version by a floating alias. Nothing errors. Over three days, refunds issued by the assistant rise 40%, because the new model is more willing to accept "my food was cold" at face value.
After the incident the team pins a dated model version; adds a 400-case eval set with refund decisions labelled by the support team, which must stay within 2 points of the previous version; tracks "refunds per 1,000 chats" as a monitored metric with an alert; and routes model upgrades through a 5% canary for 48 hours. The next upgrade shows a similar refund jump in the canary and is held back.
Follow-up questions to expect
- "How do you make LLM tests deterministic?" — Don't test the model deterministically; test your code with recorded fixtures, and test the model statistically with eval thresholds.
- "What is the most overlooked practice?" — Idempotency on side-effecting tools. Retries are everywhere in LLM stacks, and without idempotency keys they duplicate real actions.
- "How often should evals run?" — On every prompt, model, retrieval or guardrail change, and on a schedule against production samples to catch drift.