AI Safety & Guardrails

Course Content

AI Safety & Guardrails

5 sections · 50 lessons

What software engineering practices improve AI system reliability?


What you need to know

Version everything

  • Prompts, tool schemas and eval sets in git, reviewed like code.
  • Pinned model IDs with a date or version, never a "latest" alias that can change under you.
  • Index snapshots with a version, so you know which documents a given answer used.

Test like you mean it

  • Eval gate in CI: a golden dataset with thresholds; the build fails if quality drops.
  • Deterministic unit tests: record model responses as fixtures so tests of your code don't depend on the model.
  • Regression cases: every production failure becomes a test.

Handle bad outputs as normal events

  • Structured outputs with a schema; validate, and retry once with the error message if it fails.
  • Treat a malformed output as a caught error, not an exception that crashes a downstream job.

Resilience patterns

PatternWhy
TimeoutsModel calls can hang for tens of seconds
Retries with backoff and jitterHandle rate limits and transient errors
Circuit breakerStop hammering a failing provider
FallbackSmaller model, cached answer, or "try later"
Idempotency keysA retried issue_refund must not refund twice

Release safely

  • Feature flags per capability.
  • Shadow traffic: run the new version on real requests without showing users.
  • Canary: 1–5% of traffic, compare metrics, then expand.
  • One-click rollback.

Observe

  • Tracing of prompt, retrieved context, tool calls, model version, tokens, latency and cost per request (OpenTelemetry has semantic conventions for generative AI; tools like Langfuse and LangSmith build on the same idea).
  • Quality alerts: refusal rate, guardrail block rate, groundedness score, cost per request, p95 latency — not only 5xx errors.

A real-life example

A food-delivery company's support assistant is quietly switched to a newer model version by a floating alias. Nothing errors. Over three days, refunds issued by the assistant rise 40%, because the new model is more willing to accept "my food was cold" at face value.

After the incident the team pins a dated model version; adds a 400-case eval set with refund decisions labelled by the support team, which must stay within 2 points of the previous version; tracks "refunds per 1,000 chats" as a monitored metric with an alert; and routes model upgrades through a 5% canary for 48 hours. The next upgrade shows a similar refund jump in the canary and is held back.

Follow-up questions to expect

  • "How do you make LLM tests deterministic?" — Don't test the model deterministically; test your code with recorded fixtures, and test the model statistically with eval thresholds.
  • "What is the most overlooked practice?" — Idempotency on side-effecting tools. Retries are everywhere in LLM stacks, and without idempotency keys they duplicate real actions.
  • "How often should evals run?" — On every prompt, model, retrieval or guardrail change, and on a schedule against production samples to catch drift.