Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Scenario – 10: Explainability Requirement
What you need to know
The scenario: a compliance team, or a regulator, requires that users and auditors can see why the system gave each answer.
Two audiences, two surfaces
For users, in the UI
- Sources actually used, as links
- The exact passage highlighted
- A confidence band, and why it abstained
- What it could not find
For auditors, in the trace
- Query, rewritten query, chunk IDs and scores
- Index, prompt and model versions
- Guardrail verdicts and the final output
- Enough to replay the answer exactly
Users trust "I found this in section 4.2 of the November policy, and nothing about the 2023 rules" far more than an answer that simply sounds sure.
The trace record
1trace = {2 "request_id": rid, "tenant": tenant, "user": user_id_hash, "ts": now_iso(),3 "query": query, "rewritten_query": rewritten,4 "retrieved": [{"chunk_id": c.id, "doc": c.doc_id, "version": c.version,5 "retrieval_score": c.score, "rerank_score": c.rerank} for c in chunks],6 "index_version": INDEX_VERSION, "prompt_id": "policy-qa", "prompt_version": 14,7 "model": MODEL_ID, "params": {"temperature": 0.1, "max_tokens": 800},8 "output": answer, "citations": cited_ids, "guardrails": verdicts,9 "confidence": band, "latency_ms": ms, "cost_usd": cost,10}11trace_store.write(trace) # written once, by the LLM gateway, for 100% of requestsWriting it from one gateway, rather than from each feature, is what makes coverage complete.
The privacy tension
Traces hold queries and retrieved documents, so they are personal data. Encrypt them at rest, restrict access by role, set a retention period that matches your regulator's requirement, and make "delete this user" cascade into the trace store.
Metrics
- Trace completeness — share of responses with a full trace; the target is 100%.
- Audit pass rate — share of sampled traces where a reviewer can explain the answer from the record.
- Replay success — share of sampled old requests that reproduce the same sources.
A real-life example
Scenario, numbers made up. An NBFC's loan-servicing assistant is asked by an auditor to explain 50 answers from four months ago. The team has application logs, but not retrieved chunks or prompt versions, and can fully explain only 12.
They route every LLM call through one gateway that writes the trace, version prompts and index snapshots, and show citations with highlighted passages in the agent's UI. At the next audit, all 50 sampled answers are explained from the stored trace, and 48 replay with the same sources (the other two used documents since deleted under the retention policy, which is documented). Agents also start clicking citations before relaying answers to customers.
Follow-up questions to expect
- "Isn't storing everything expensive?" — Traces are text and compress well; tier old traces to cheap storage and keep a searchable index of IDs and metadata.
- "How do you explain why a chunk was chosen?" — Show its retrieval and rerank scores, matched terms, and filters applied; do not ask the model to invent a reason afterwards.
- "What if a document changed since the answer?" — Store the document version with the trace, so the auditor sees what the system saw at the time.