Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

A regulatory audit requires you to explain every AI-generated answer from the past 6 months — documents retrieved, model used, prompt sent, response returned. You never built logging. How do you design a retroactive audit trail for LLM outputs?


Every call, one tamper-evident recordAll model callspass one gatewayRecord: chunks,versions, outputChain eachrecord tothe last hashWrite-oncestorage,set retentionPast retrieval could not be recovered, and the report said so.
Logging at one gateway is what makes coverage complete; logging at each call site depends on every team remembering.

What you need to know

What can be recovered

ItemWhere it may surviveConfidence
ResponsesChat history tableHigh
Model and versionDeploy history, config in git, provider invoicesMedium to high
Prompt templateGit history joined to deploy timestampsMedium
Request metadataApplication logs, if retention covers 6 monthsVaries
Retrieved documentsNowhere, unless chunk IDs were loggedUsually none

Re-running retrieval today against a changed index tells you what the system would retrieve now, not what it did then. Label anything like that clearly as a reconstruction, never as the original record.

The gap analysis

  1. Recoverable with high confidence — what, from where, how verified.
  2. Inferred — what, and the method (for example, prompt version inferred from deploy time).
  3. Unrecoverable — stated plainly.
  4. Remediation — what is being built, and by when.
  5. Legal review — of the wording before it goes to the regulator.

Make it structural

Each audit record holds: request ID, tenant and user, timestamp, raw and rewritten query, retrieved chunk IDs with scores and index version, prompt template ID and version plus a hash of the rendered prompt, model and parameters, the raw output, guardrail verdicts, latency, cost and any feedback.

Python
import hashlib, jsondef append_audit(record: dict, prev_hash: str) -> str:    record["prev_hash"] = prev_hash                          # chains each record to the one before    body = json.dumps(record, sort_keys=True, ensure_ascii=False)    record_hash = hashlib.sha256(body.encode()).hexdigest()    worm_store.put(f"audit/{record['ts'][:10]}/{record['request_id']}.json",                   body, object_lock_days=RETENTION_DAYS)   # write-once storage    return record_hash

Write-once storage (such as S3 Object Lock in compliance mode) stops records being changed or deleted during retention, and hash chaining makes any tampering detectable. Retention length is set by your regulator.

The privacy side

Audit records contain personal data. Encrypt them, restrict access by role, and document the retention policy with the audit design. Where the law requires deleting personal data before audit retention ends, store personal fields separately (or encrypted with a per-user key you can destroy) so the rest of the record survives.

A real-life example

Scenario, numbers made up. An insurance company's claims assistant has answered about 400,000 questions in six months when the regulator asks for a full explanation of 200 sampled answers. There is no audit logging.

The team recovers all 200 responses from the chat database, identifies the model for 196 from deploy logs and invoices, and matches prompt versions for 181 from git and deploy times. Retrieved documents cannot be recovered for any of them; the report says so, and describes what the index contained at the time from document management records. The regulator accepts the gap analysis with a condition: the gateway audit log goes live within 30 days. It ships in 9 days, and the next sample of 200 is fully explained from stored records.

Follow-up questions to expect

  • "Why not log at each call site?" — Coverage depends on every team remembering; one new feature without logging breaks the audit. A gateway makes it automatic.
  • "Why store a hash of the rendered prompt?" — It proves exactly what was sent without keeping another full copy, and you can store the template and variables to rebuild it.
  • "How do you prove logs weren't altered?" — Write-once storage plus hash chaining; any change breaks the chain.