Course Content
Live Coding Interview Prep
7 sections · 50 lessons
Build output guardrails for filtering unsafe or irrelevant responses.
What you need to know
Input checks guard what goes into the model; output guardrails guard what comes out. Typical checks and their actions:
| check | example | action |
|---|---|---|
| Empty or truncated | whitespace only | regenerate |
| Secret leakage | an API key from the context | block |
| Personal data | emails, card numbers | redact |
| Format | invalid JSON for a structured endpoint | regenerate |
| Policy | unsafe or off-topic content | block with a safe message |
| Grounding | claims not supported by the sources | regenerate or add a warning |
Fail open or fail closed? If the policy classifier times out, does the answer go out unchecked (open) or not at all (closed)? A children's education app should fail closed; an internal coding assistant might fail open. The answer must be a decision, not an accident.
Luhn check. Card numbers carry a check digit. Running the Luhn algorithm on a 13–19 digit match separates real card numbers from order ids and phone numbers, which cuts false redactions sharply.
1import json, re2from collections.abc import Callable3from dataclasses import dataclass, field4import jsonschema56EMAIL = re.compile(r"\b[\w.+-]+@[\w-]+\.[\w.]{2,}\b")7CARD = re.compile(r"\b(?:\d[ -]?){12,18}\d\b")8SECRET = re.compile(r"\b(sk-[A-Za-z0-9_-]{16,}|AKIA[0-9A-Z]{16})\b")9FENCE = re.compile(r"^`{3}(?:json)?\s*|\s*`{3}$") # a Markdown code fence1011def luhn_ok(number: str) -> bool:12 digits = [int(d) for d in re.sub(r"\D", "", number)][::-1]13 total = sum(d if i % 2 == 0 else (d * 2 - 9 if d > 4 else d * 2) for i, d in enumerate(digits))14 return total % 10 == 01516def redact(text: str) -> str:17 text = EMAIL.sub("[email]", text)18 return CARD.sub(lambda m: "[card]" if luhn_ok(m.group()) else m.group(), text)1920@dataclass21class Verdict:22 action: str # allow | redact | regenerate | block23 text: str24 reasons: list[str] = field(default_factory=list)2526def check_output(text: str, schema: dict | None = None,27 classify: Callable[[str], str] | None = None) -> Verdict:28 if not text.strip():29 return Verdict("regenerate", text, ["empty response"])30 if SECRET.search(text):31 return Verdict("block", "", ["credential in output"])32 cleaned = redact(text)33 reasons = ["personal data redacted"] if cleaned != text else []34 if schema is not None:35 try:36 jsonschema.validate(json.loads(FENCE.sub("", cleaned.strip())), schema)37 except (json.JSONDecodeError, jsonschema.ValidationError) as exc:38 detail = getattr(exc, "message", None) or str(exc) # one-line reason39 return Verdict("regenerate", cleaned, reasons + [f"schema: {detail[:60]}"])40 if classify is not None and classify(cleaned) != "safe":41 return Verdict("block", "", reasons + ["policy violation"])42 return Verdict("redact" if reasons else "allow", cleaned, reasons)The tricky parts:
- Secrets block, personal data redacts. A leaked key means something upstream is badly wrong, so nothing goes out; an email address can be masked and the rest of the answer is still useful.
FENCE.substrips a Markdown code fence (three backticks, optionally followed byjson), which models add often and which makesjson.loadsfail before the schema is even checked.- The Luhn lambda only replaces matches that pass the checksum, so "order 4000123412341235" (fails Luhn) survives while a real card number is masked.
- Order matters. Cheap checks first, and each block or regenerate returns immediately, so the classifier call only happens for outputs that passed everything else.
Complexity: each regex pass is O(L) for output length L; Luhn is O(digits) per match; schema validation is O(size of the JSON); the classifier is one model call. Space O(L).
A real-life example
1schema = {"type": "object", "required": ["refund_days"],2 "properties": {"refund_days": {"type": "integer"}}}3fence = "`" * 34outputs = ["Mail us at help@shop.in. Card on file: 4111 1111 1111 1111.",5 "Your order 4000123412341235 is packed.",6 "Use key sk-live_ab12cd34ef56gh78ij to call the API.",7 fence + 'json\n{"refund_days": 5}\n' + fence,8 '{"refund_days": "five"}',9 " "]10for out in outputs:11 structured = out.lstrip().startswith(("{", fence))12 v = check_output(out, schema=schema if structured else None)13 print(v.action, v.reasons)14# redact ['personal data redacted']15# allow []16# block ['credential in output']17# allow []18# regenerate ["schema: 'five' is not of type 'integer'"]19# regenerate ['empty response']20print(check_output(outputs[0]).text)21# Mail us at [email]. Card on file: [card].| output | check that decided | action |
|---|---|---|
| email and a real test card number | redaction (card passes Luhn) | redact |
| 16-digit order id | Luhn fails, so not a card | allow |
| API key | secret pattern | block |
| JSON in a fence | fence stripped, schema passes | allow |
"five" instead of 5 | schema | regenerate |
| whitespace | empty check | regenerate |
Payment and banking assistants run exactly this kind of pipeline, because one leaked card number in a chat transcript is a compliance incident.
Follow-up questions to expect
- "How do you guard a streaming response?" — You cannot un-send tokens. Buffer until a sentence ends and check each sentence, or hold back the first few hundred milliseconds, or accept post-hoc correction. Say which you chose and why.
- "How do you check grounding?" — Per claim, ask an NLI model or a small LLM whether the retrieved context supports it; word-overlap scores are only a rough first version.
- "A regenerate loop never converges?" — Allow one regeneration, then fall back to a safe canned answer and log the case for review.