LLMOps & Deployment

Course Content

LLMOps & Deployment

6 sections · 40 lessons

What are guardrails, and how do they ensure safe model behavior in production?


What you need to know

Where guardrails sit

  1. Input checks — length and attachment limits, PII detection and redaction (for example Microsoft Presidio or regex for card numbers), topic scope, prompt-injection and jailbreak classifiers.
  2. Model call — with a system prompt that states the rules; helpful, but not a guardrail on its own.
  3. Tool-call checks — the agent may call only allow-listed tools, with arguments validated against a schema and business limits (for example refund amount at most Rs 2,000 without a human).
  4. Output checks — JSON schema validation, PII and toxicity scans, groundedness against retrieved context, banned phrases (legal promises, competitor names).
  5. Action on failure — redact, retry once, replace with a safe canned answer, or hand off to a human.

Tools

Plain regex and Pydantic validation cover a lot. For classifiers: Llama Guard (an open safety classifier model), provider moderation endpoints, and frameworks such as NeMo Guardrails and Guardrails AI that wire checks around calls.

The latency and error budget

Each classifier adds time — often tens to a few hundred milliseconds — and makes mistakes. Two design choices keep this cheap:

  • Order by cost. Regex and schema checks take about a millisecond; run them always. Run model-based checks only where risk is high.
  • Run in parallel. Start the input classifier and the main LLM call at the same time; if the classifier flags the input, discard the answer before it is shown.

Decide fail closed (block when the check fails or times out) for high-severity rules like data leakage, and fail open (allow and log) for low-severity ones like tone.

Streaming makes output checks harder

If tokens stream to the user, a full-output check comes too late. Common fixes: check each sentence before releasing it, or buffer structured outputs completely and stream only plain prose.

A real-life example

A fintech support bot gets this message: "My card 4111 1111 1111 1111 was charged twice, OTP was 482913, please refund now."

  • The input guardrail redacts the card number and OTP before the text reaches the model or the logs (a regex plus a Luhn check, about 1 ms).
  • The model drafts: "I have processed your refund of Rs 12,000." The tool-call guardrail sees that the refund tool was not called with an approved ticket, and the output check blocks the phrase "processed your refund".
  • The user gets a safe template: "I have raised a dispute; you will hear back within 3 working days," and the case goes to a human agent.

Over a month, the team reviews 200 logged blocks and finds 14% were false positives on the phrase check, so they narrow the rule.

Follow-up questions to expect

  • "Isn't a strong system prompt enough?" — No. Prompt injection can override instructions; guardrails are separate code or models that the user's text cannot rewrite.
  • "How do you stop prompt injection from retrieved documents?" — Treat retrieved text as data, never instructions; limit what tools the model can call; validate every tool argument; require human approval for risky actions.
  • "How do you know if guardrails are too strict?" — Log every block with the trace ID, review a sample, and track the false-positive rate as a metric.