Live Coding Interview Prep

Course Content

Live Coding Interview Prep

7 sections · 50 lessons

Build output guardrails for filtering unsafe or irrelevant responses.


What you need to know

Input checks guard what goes into the model; output guardrails guard what comes out. Typical checks and their actions:

checkexampleaction
Empty or truncatedwhitespace onlyregenerate
Secret leakagean API key from the contextblock
Personal dataemails, card numbersredact
Formatinvalid JSON for a structured endpointregenerate
Policyunsafe or off-topic contentblock with a safe message
Groundingclaims not supported by the sourcesregenerate or add a warning

Fail open or fail closed? If the policy classifier times out, does the answer go out unchecked (open) or not at all (closed)? A children's education app should fail closed; an internal coding assistant might fail open. The answer must be a decision, not an accident.

Luhn check. Card numbers carry a check digit. Running the Luhn algorithm on a 13–19 digit match separates real card numbers from order ids and phone numbers, which cuts false redactions sharply.

Python
import json, refrom collections.abc import Callablefrom dataclasses import dataclass, fieldimport jsonschemaEMAIL = re.compile(r"\b[\w.+-]+@[\w-]+\.[\w.]{2,}\b")CARD = re.compile(r"\b(?:\d[ -]?){12,18}\d\b")SECRET = re.compile(r"\b(sk-[A-Za-z0-9_-]{16,}|AKIA[0-9A-Z]{16})\b")FENCE = re.compile(r"^`{3}(?:json)?\s*|\s*`{3}$")     # a Markdown code fencedef luhn_ok(number: str) -> bool:    digits = [int(d) for d in re.sub(r"\D", "", number)][::-1]    total = sum(d if i % 2 == 0 else (d * 2 - 9 if d > 4 else d * 2) for i, d in enumerate(digits))    return total % 10 == 0def redact(text: str) -> str:    text = EMAIL.sub("[email]", text)    return CARD.sub(lambda m: "[card]" if luhn_ok(m.group()) else m.group(), text)@dataclassclass Verdict:    action: str                                   # allow | redact | regenerate | block    text: str    reasons: list[str] = field(default_factory=list)def check_output(text: str, schema: dict | None = None,                 classify: Callable[[str], str] | None = None) -> Verdict:    if not text.strip():        return Verdict("regenerate", text, ["empty response"])    if SECRET.search(text):        return Verdict("block", "", ["credential in output"])    cleaned = redact(text)    reasons = ["personal data redacted"] if cleaned != text else []    if schema is not None:        try:            jsonschema.validate(json.loads(FENCE.sub("", cleaned.strip())), schema)        except (json.JSONDecodeError, jsonschema.ValidationError) as exc:            detail = getattr(exc, "message", None) or str(exc)       # one-line reason            return Verdict("regenerate", cleaned, reasons + [f"schema: {detail[:60]}"])    if classify is not None and classify(cleaned) != "safe":        return Verdict("block", "", reasons + ["policy violation"])    return Verdict("redact" if reasons else "allow", cleaned, reasons)

The tricky parts:

  • Secrets block, personal data redacts. A leaked key means something upstream is badly wrong, so nothing goes out; an email address can be masked and the rest of the answer is still useful.
  • FENCE.sub strips a Markdown code fence (three backticks, optionally followed by json), which models add often and which makes json.loads fail before the schema is even checked.
  • The Luhn lambda only replaces matches that pass the checksum, so "order 4000123412341235" (fails Luhn) survives while a real card number is masked.
  • Order matters. Cheap checks first, and each block or regenerate returns immediately, so the classifier call only happens for outputs that passed everything else.

Complexity: each regex pass is O(L) for output length L; Luhn is O(digits) per match; schema validation is O(size of the JSON); the classifier is one model call. Space O(L).

A real-life example

Python
schema = {"type": "object", "required": ["refund_days"],          "properties": {"refund_days": {"type": "integer"}}}fence = "`" * 3outputs = ["Mail us at help@shop.in. Card on file: 4111 1111 1111 1111.",           "Your order 4000123412341235 is packed.",           "Use key sk-live_ab12cd34ef56gh78ij to call the API.",           fence + 'json\n{"refund_days": 5}\n' + fence,           '{"refund_days": "five"}',           "   "]for out in outputs:    structured = out.lstrip().startswith(("{", fence))    v = check_output(out, schema=schema if structured else None)    print(v.action, v.reasons)# redact ['personal data redacted']# allow []# block ['credential in output']# allow []# regenerate ["schema: 'five' is not of type 'integer'"]# regenerate ['empty response']print(check_output(outputs[0]).text)# Mail us at [email]. Card on file: [card].
outputcheck that decidedaction
email and a real test card numberredaction (card passes Luhn)redact
16-digit order idLuhn fails, so not a cardallow
API keysecret patternblock
JSON in a fencefence stripped, schema passesallow
"five" instead of 5schemaregenerate
whitespaceempty checkregenerate

Payment and banking assistants run exactly this kind of pipeline, because one leaked card number in a chat transcript is a compliance incident.

Follow-up questions to expect

  • "How do you guard a streaming response?" — You cannot un-send tokens. Buffer until a sentence ends and check each sentence, or hold back the first few hundred milliseconds, or accept post-hoc correction. Say which you chose and why.
  • "How do you check grounding?" — Per claim, ask an NLI model or a small LLM whether the retrieved context supports it; word-overlap scores are only a rough first version.
  • "A regenerate loop never converges?" — Allow one regeneration, then fall back to a safe canned answer and log the case for review.