- MantraMindAI
- Blog
- Responsible AI & Security
AI governance that actually runs: gates, registries, audit trails
Jai Rao
August 22, 202616 min read
Most AI policies are wishes because nothing checks them. What an enforced control looks like: registry admission, blocking eval gates, schema checks, decision records.
A support ticket arrives asking why your system declined someone in March. Answering it takes four facts: which model version ran, what data it saw, which prompt and policy version were in force, and who approved putting that version in front of customers. If those four facts live in four different places — a container tag, a Slack thread, a spreadsheet maintained by someone who left, and nobody's memory — you do not have an abstract governance gap. You have a question you cannot answer, about a decision you already made.
The usual response to this is a document. Someone writes a responsible-AI policy, it says models must be evaluated before release and high-risk systems require human review, it gets approved, and nothing in the delivery path changes. Six months later the same ticket is still unanswerable. The reason is simple enough to state as a rule: a policy that is not enforced by a pipeline is not a control, it is a wish. What follows is about the enforced version — the registry that refuses a model, the CI job that blocks a deploy, the schema that rejects an output, and the records that make the March question a ten-minute lookup.
What makes something a control rather than a wish
A control has four properties, and if any one is missing you have documentation instead. It has a trigger: a specific event in the delivery path where it runs — a merge, a registry admission, a promotion to production, an inference call. It has a mechanism that can say no, meaning it returns a non-zero exit code or throws, rather than printing a warning into a log nobody reads. It fails closed: if the eval service is unreachable, the deploy stops, because "we could not check" is not the same as "it passed". And it leaves a record, so that later you can show not only that the rule existed but that it ran on this change.
Here is the same set of intentions written both ways. The left column is what most policies contain. The right column is what you can actually point at.
| Written as policy | Enforced as a control |
|---|---|
| Teams should document intended use. | Registry admission rejects a model version whose card is missing intended_use or out_of_scope. |
| Model quality must not regress. | CI runs the versioned eval suite and exits non-zero when any tracked slice drops more than its declared tolerance. |
| High-risk systems require human review. | Promotion for tier-1 systems reads an approval record signed by the named owner; the job blocks when it is absent or older than the artefact. |
| We monitor for bias. | A scheduled job computes per-group metrics on production traffic and opens an incident with an owner and a due date when a threshold is breached. |
| Outputs must be validated. | Every model response is parsed against a schema at the boundary; failures escalate to a deterministic path and increment a rejection counter. |
Exceptions are not the enemy of this — pretending you will never need one is. Build a break-glass path deliberately: it requires a reason string, it names a person, it writes to the same audit log as everything else, and it expires. A bypass that silently persists is how a gate becomes decorative. A bypass that shows up on a weekly list with an owner and an expiry date is a control with a documented escape hatch, which is what mature systems have.
Tiering by consequence and reversibility
The fastest way to destroy the credibility of a governance program is to apply the same load to a spam classifier and a lending decision. When a spam filter is wrong, someone misses an email and finds it in a folder; the cost is small and the user can undo it themselves. When a lending model is wrong, a person loses access to credit, and undoing it requires them to know an appeal exists, find it, and win. Requiring the same review board, the same sign-off, and the same documentation for both teaches engineers that the checklist is theatre. Once they believe that, they route around it, and your controls now cover only the teams that were never the risk.
Two questions do most of the tiering work. How bad is a wrong output for the person on the other end? And how hard is it to reverse? Add a third when the system acts rather than advises, and a fourth for scale — a hundred decisions a day and ten million a day are different systems even with identical code.
| Tier | Shape | Required artefacts | Gates |
|---|---|---|---|
| 3 | Internal, a human reads every output before anything happens, trivially reversible | Card, named owner, eval set | CI eval gate |
| 2 | Customer-facing, affects experience or small amounts of money, reversible with effort | Plus data provenance, per-slice evals, monitoring with alerts | CI eval gate, owner sign-off on promotion |
| 1 | Touches money, employment, health, safety or legal standing; hard to reverse; may act without a human | Plus documented review path, appeal route, retained decision records, scheduled re-review | All of the above, plus recorded pre-deploy approval and a retirement owner |
One detail matters more than the tier definitions: the tier is a property of the deployment, not the model. The same general-purpose model summarising internal meeting notes and drafting patient-facing instructions is a tier-3 system and a tier-1 system. Tier the use case, record the tier in the registry entry, and let the gates read it from there. Otherwise every conversation becomes an argument about how risky the model is in the abstract, which nobody can settle.
The artefacts that carry weight
Five things end up doing real work. A system card that names the owner, the tier, what the thing is for, and what it must not be used for. An intended-use and out-of-scope list, which is the part of the card people skip and the part that settles the most arguments. A data provenance record: where the training or retrieval corpus came from, under what permission, and when it was last refreshed. A versioned eval set that lives in the repository and changes through review like code. And an incident log where degradations and near-misses get written down with what changed as a result.
What turns these from documentation into artefacts is location and enforcement. A card in a wiki is a page. A card in the repo, versioned with the code, parsed by a gate that refuses to promote a model when a required field is empty, is a control surface. The fragment below is the minimum that has ever been useful to me — note that out_of_scope is written as concrete prohibitions rather than platitudes.
system: loan_prefill_assistantversion: 7tier: 1owner_email: priya.raman@example.comdeputy_email: sam.okafor@example.comintended_use: | Pre-fills application fields from documents the applicant uploaded, for review by a human underwriter before any decision is recorded.out_of_scope: - producing an approve/decline recommendation - reading documents the applicant did not upload themselves - any path where the output is stored without underwriter confirmationdata_provenance: corpus: internal_scanned_forms_2019_2024 permission_basis: applicant_consent_v3 last_refreshed: 2026-05-02evaluation: suite: evals/loan_prefill/v7 run_id: er_9f31c2 slices: [handwritten, low_resolution, non_english, redacted]retention: decision_records_days: 2555The card is not the interesting part on its own. Its value comes entirely from the next section: something has to read it and refuse.
The registry as the chokepoint
If you build one piece of governance infrastructure, build a model registry that will not accept an entry without an owner and a resolvable eval run. Everything else can hang off it. The registry is where "which version is in production" stops being folklore, and it is the natural place to enforce the artefacts, because promotion is a single narrow event you control.
This is an admission check, roughly as I would write it for a first pass. Watch the digest comparison in the middle.
REQUIRED = ("owner_email", "tier", "card_path", "eval_run_id", "model_digest")class Reject(Exception): passdef admit(entry, eval_store, directory): missing = [k for k in REQUIRED if not entry.get(k)] if missing: raise Reject(f"missing required fields: {missing}") if not directory.is_active(entry["owner_email"]): raise Reject(f"owner {entry['owner_email']} is not an active account") run = eval_store.get(entry["eval_run_id"]) if run is None: raise Reject("eval_run_id does not resolve to a stored run") if run["model_digest"] != entry["model_digest"]: raise Reject("eval was run against a different artefact") if entry["tier"] == 1 and not entry.get("approval_record_id"): raise Reject("tier 1 promotion requires a recorded approval") return entryThe digest check is the line that earns its keep. An eval run attached to a different build is worse than having no eval at all, because it looks like evidence. That mismatch happens constantly in real pipelines — someone reruns the training job, the metrics from the previous run are still sitting in the dashboard, and the number gets copied forward. Comparing the hash of the artefact you evaluated to the hash of the artefact you are promoting costs four lines and catches it every time.
Two more properties worth designing in from the start. Registry entries are append-only: you supersede a version, you never edit one, because an audit trail that points at a mutable row proves nothing. And the entry stores references — card path at a commit, eval run id, dataset snapshot id — rather than copies, so the record stays small while remaining resolvable years later.
Blocking the deploy when the eval suite regresses
The eval gate is the control that most changes team behaviour, because it moves quality from something you assert in a review to something the build enforces. The mechanics are unglamorous: run the versioned suite, compare against a pinned baseline, and exit non-zero on a real drop.
Two design choices decide whether this works. Compare per slice, not on an aggregate — an average can hold steady while performance on non-English inputs collapses, and the aggregate is exactly where that hides. And set tolerances from measured run-to-run variance rather than intuition, otherwise you either block on noise or wave through genuine losses.
import json, sysbase = json.load(open("evals/baseline.json")) # pinned, reviewed, in gitnew = json.load(open("evals/latest.json"))tol = json.load(open("evals/tolerance.json")) # per slice, from measured variancefailures = []for name, baseline_score in base["slices"].items(): if name not in new["slices"]: failures.append(f"{name}: slice missing from this run") continue drop = baseline_score - new["slices"][name] if drop > tol.get(name, 0.01): failures.append(f"{name}: {baseline_score:.3f} -> {new['slices'][name]:.3f}")for f in failures: print("REGRESSION", f)sys.exit(1 if failures else 0)The clause that surprises people is treating a missing slice as a failure. Deleting the slice that fails is the most common way an eval gate gets defeated, and it rarely happens maliciously — someone is under deadline pressure, the non-English slice is red, removing it makes the build green, and the commit message says "clean up flaky evals". Requiring every baseline slice to be present in every run makes that a visible act rather than a quiet one.
Pair it with a rule about the baseline file: changing it needs review by the system owner, and the pull request has to say why the bar moved. Improving a number by lowering the threshold is a legitimate thing to do occasionally and a catastrophic thing to do invisibly. For generative outputs, pin what you can — temperature, seed, model version, prompt version — and accept that some suites need to be graded as a distribution over repeated runs rather than a single score. Measure the variance first; the tolerance file should be an empirical artefact.
Schema gates at the output boundary
Everything so far runs before deploy. At runtime you need one more control, and it is the cheapest of the lot: nothing a model produces enters your system unparsed. Treat the output as a proposal from an untrusted source, validate its shape, and have a defined behaviour for failure.
This gate does three things: it parses, it checks the claimed citation against documents the request was actually allowed to see, and it escalates rather than improvising when either fails.
from typing import Literalfrom pydantic import BaseModel, Field, ValidationErrorclass Triage(BaseModel): category: Literal["billing", "outage", "how_to", "other"] confidence: float = Field(ge=0.0, le=1.0) evidence_doc_id: str = Field(min_length=1)def gate(raw: str, allowed_docs: set[str]): try: out = Triage.model_validate_json(raw) except ValidationError as err: metrics.incr("triage.schema_reject") return Escalate("unparseable", detail=str(err)) if out.evidence_doc_id not in allowed_docs: metrics.incr("triage.citation_outside_scope") return Escalate("bad_citation", detail=out.evidence_doc_id) if out.confidence < 0.75: metrics.incr("triage.low_confidence") return Escalate("needs_human") return outThe rejection counters are the part people forget to wire up, and they are the highest-value signal in the whole system. A schema rejection rate that jumps from a fraction of a percent to five percent overnight means something upstream moved — a provider updated a model, a prompt was edited, a retrieval index rebuilt. That is your earliest warning, and it arrives before any user complains.
Say the honest thing about the limit, though. A schema gate checks shape and attributability, not truth. It will happily pass a confidently wrong category carrying a real document id. It stops malformed and unattributable output; catching wrong-but-well-formed output is what the eval suite, the per-group monitoring, and the human review path are for. Anyone who tells you validation makes a model safe is selling something.
Answering the question six months later
Back to the March ticket. Being able to answer it means one append-only record per consequential decision, written at the time, containing pointers rather than a copy of everything. The fields below are the ones I have never regretted storing.
{ "decision_id": "d_01JX9K2M", "recorded_at": "2026-03-14T09:22:41Z", "subject_ref": "sha256:7c1f...", "system": "loan_prefill_assistant", "registry_version": 7, "model_digest": "sha256:9ab3...", "prompt_version": "pf_2026_02_11", "policy_version": "pol_14", "retrieval_snapshot": "idx_2026_03_12", "gate_result": "passed", "confidence": 0.88, "human_action": {"actor": "underwriter_412", "outcome": "confirmed"}, "approval_record_id": "ap_3391"}Note what is not in there: no raw application text, no name, no document contents. The subject is a salted hash you can rejoin through a separately access-controlled mapping, and the inputs are referenced by snapshot id. That keeps the log cheap to retain for years and keeps a routine debugging query from being a data-protection event. The retention window belongs in the card, and it should be a decision someone made, not whatever the default log rotation happens to be.
The record only works if its references still resolve, which is why immutable registry entries matter. registry_version: 7 is worthless if the row for version 7 has been edited twice since March. Test this like any other capability: pick a real decision from a quarter ago and try to reconstruct it end to end. If it takes a week of archaeology, the trail is decorative.
It is worth noting that regulatory regimes in several jurisdictions are moving in this direction — expecting organisations to be able to describe what a system does, show that it was tested, and explain a decision that affected someone. The artefacts above overlap heavily with what those regimes tend to ask for, which is a good argument for building them for engineering reasons first, since you get an operable system either way. What specifically applies to your product, sector, and geography is a question for your counsel; nothing here is legal advice.
Ownership without inventing an org chart
"Everyone is responsible for responsible AI" is functionally identical to nobody being responsible, and it is the single most common failure in governance documents. The fix is not a new department. It is filling a small number of named slots, and storing the names somewhere a gate can read them.
- One accountable owner per system. A person, with an email in the registry entry, who can be paged, who signs tier-1 promotions, and who is answerable when the thing misbehaves. Not a team alias — aliases cannot be accountable.
- A deputy. People take holidays. A control that stalls for two weeks gets bypassed.
- An engineering owner for the gates themselves. The pipeline, the registry, and the eval harness are a product with users, and unowned tooling rots into "warn" mode.
- A reviewer who is not the author for baseline changes, tier assignments, and tier-1 approvals. This is the only real defence against the bar quietly moving.
- An owner for the incident log — meaning someone runs the review and makes sure each entry produces a change, not someone who personally fixes everything.
- A retirement owner. Models have to be turned off. The most common orphan I see is a model still serving live traffic that nobody has claimed for a year.
Make the owner field verifiable rather than aspirational: a scheduled job that checks every registry owner still resolves to an active account, and files a ticket when one does not, catches organisational drift better than any annual review. A central board is genuinely useful for exactly two jobs — settling tier assignments and adjudicating exceptions. Route every change through it and you have rebuilt the committee document, in meeting form.
How to tell whether any of this is real
There is a quick diagnostic. Pick a consequential decision your system made three months ago and answer three questions: which model version produced it, on what data, approved by whom. If you can do that in ten minutes from records rather than memory, the audit trail works. Then pick your most recent model deploy and ask what would have stopped it if the eval suite had dropped fifteen points on one slice. If the answer is "someone would probably have noticed", you have a wish.
The failure patterns are consistent enough to check for directly. Grep your CI configuration for continue-on-error and its equivalents; gates parked in warning mode for months are the most common form of governance theatre. Look at whether your eval suite has changed since launch — a set frozen at launch is a set you have overfitted to, and the cheap fix is a standing rule that every incident contributes at least one case. Count how many systems are tier 1; if it is most of them, tier inflation has already happened and engineers are working around the process. Check whether any break-glass bypass is still active past its expiry. And look at where incidents are recorded: if the answer is a Slack channel, they are not recorded, because nothing there can be counted or reviewed.
If you are starting from nothing, the order that works is smallest-blast-radius first. Stand up a registry that requires an owner and a resolvable eval run. Put one eval suite in CI as a blocking gate for one system. Write one card format and make admission refuse an incomplete one. Add decision records for your single highest-consequence system. Then, once those exist, write the tiering policy — because at that point the policy is describing controls you already have, rather than describing controls somebody else will need to build. That sequence is a couple of weeks of engineering work. The committee version takes a quarter and enforces nothing.