AI Safety & Guardrails

Course Content

AI Safety & Guardrails

5 sections · 50 lessons

What is the checkpoint and rollback pattern for AI systems?


One release bundle, reverted as a unitModel: pinneddated versionPrompt:support_v42Index:immutable snapshotGuardrailpolicy_v17Tool policy_v9topbottomFlipping the bundle alias restored prompt and index together in 90 seconds.
Rolling back only the prompt lands you in a combination nobody tested; the bundle is the unit of release and of rollback.

What you need to know

Release-level checkpoints

A release bundle is a manifest listing:

Text
model:        vendor-model-2026-05-14prompt:       support_v42index:        kb_snapshot_2026_09_20guardrails:   policy_v17tools:        tool_policy_v9

Deploy and roll back the whole bundle. Rolling back only the prompt while keeping the new index creates a combination nobody ever tested.

Practices that make this real:

  • Immutable index snapshots with blue-green alias switching: build the new index beside the old one, point the alias at it, and point it back to roll back.
  • Backward-compatible migrations, so a rollback is not blocked by a data change you cannot undo.
  • One toggle for on-call to execute, without a code deploy.

Run-level checkpoints (agents)

An agent saves its state after every step:

  • Resume after a crash from the last good step, instead of restarting and repeating side effects.
  • Human-in-the-loop: pause before a sensitive step, let a person inspect and edit the state, then continue. LangGraph supports this with checkpointers and interrupts.
  • Time travel: go back to an earlier checkpoint and try a different branch.

Side effects are the hard part

Weights and configs roll back; sent emails and issued refunds do not. For every side-effecting tool:

  • Idempotency key: the same request twice has one effect.
  • Dry-run mode: see what would happen.
  • Compensating action: a documented way to undo (reverse a refund, send a correction).
  • Outbox: record intended actions before executing, so you know exactly what went out.

A real-life example

A bank's customer chatbot releases a new prompt and a re-embedded knowledge base together on a Friday. By Saturday morning, answers about home-loan documents are citing a deleted 2023 policy. On-call switches the bundle alias back to the previous release in 90 seconds, restoring both the old prompt and the old index snapshot.

The root cause, found on Monday: the re-embedding job included an archive folder. Because the old index was an immutable snapshot, rollback was instant. In the previous year, before bundles, a similar incident took four hours, because the team rolled back the prompt, found it didn't help, and had to rebuild the old index from scratch.

Follow-up questions to expect

  • "How fast should rollback be?" — Minutes, by one person on call, without a deploy — and proven in a drill.
  • "How long do you keep old snapshots?" — At least until the new release has been stable for a defined period, and longer if you need them for audits or appeals.
  • "What if the vendor retires the old model?" — That is why you track deprecation dates and keep a tested fallback model in the bundle options.