Course Content
Enterprise AI Solutions Architecture
13 sections · 29 lessons
Failure Semantics: Idempotency, Timeouts and Compensation
In the second week of the pilot, a hardship specialist clicked "Draft letter", waited, and clicked again. Meanwhile the gateway timed out the first request and retried it. DocGen received three requests and created three drafts for the same case. One of them had been generated before the specialist changed the plan length from six months to nine, so it carried the old figures. It was the one she opened and approved. QA caught it in the 5% sample, two days after it was sent.
None of this was about AI. It was a classic distributed systems failure: no timeouts that fitted together, no idempotency and no rule about which draft was current. But AI makes these failures more likely, because model calls are slow. A call that takes 12 seconds invites double-clicks, hits timeouts and triggers retries far more often than a call that takes 50 milliseconds.
This lesson defines Meridian's failure semantics: what every interface does when it is slow, down or called twice. It produces part B of MER-07.
Timeouts that add up
A timeout is not the same as a p95. The p95 describes normal behaviour; a timeout says when to give up. Set timeouts from two directions: long enough that normal slow calls finish, and short enough that the sum of sequential steps fits the user-facing deadline.
The letter draft has a 25-second p95 requirement. Here is how its budget is divided.
| Step | Runs | p95 | Timeout |
|---|---|---|---|
| Calculator, account facts, notes | In parallel | 450 ms | 1.5 s |
| Model draft, about 900 tokens | Sequential | 11 s | 15 s |
| Check pass: figures, paragraphs, claims | Sequential | 4 s | 6 s |
| DocGen create draft | Sequential | 300 ms | 2 s |
| Worst case before giving up | 24.5 s |
The worst case, with every step at its timeout, still fits under 25 seconds. The design also passes a deadline down the chain: each step receives the time remaining and uses the smaller of that and its own timeout. If the model takes 14 seconds, the check pass gets only what is left, and the request fails cleanly rather than running over.
Idempotency
An operation is idempotent if doing it twice has the same effect as doing it once. Reads are naturally idempotent. Draft creation is not, unless you make it so. The standard technique is an idempotency key: a value derived from what makes the request unique, sent with the request, so the receiver can recognise a repeat.
1import hashlib2import json34def draft_key(case_id: str, plan: dict, calc_version: str, template_id: str) -> str:5 """Same case, plan, figures and template means the same draft."""6 material = json.dumps(7 {"case": case_id, "plan": plan, "calc": calc_version, "tpl": template_id},8 sort_keys=True,9 )10 return hashlib.sha256(material.encode()).hexdigest()[:32]1112def create_draft_once(docgen, ctx, plan: dict, calc: dict, template_id: str, body: str) -> str:13 key = draft_key(ctx.case_id, plan, calc["version"], template_id)14 draft_id = docgen.create_draft(15 case_id=ctx.case_id, template_id=template_id, body=body,16 figures=calc["figures"], idempotency_key=key, token=ctx.obo_token,17 )18 docgen.supersede_others(case_id=ctx.case_id, keep=draft_id, token=ctx.obo_token)19 return draft_idThe key is built from the case, the plan terms, the calculator result version and the template. It deliberately leaves out the letter body, because the model writes different words each time; a retry must not count as a new draft just because the wording changed. DocGen stores the key and, if it sees the same key again, returns the existing draft instead of creating another. That server-side check is the real guarantee; a check only in the caller would still race when two requests arrive together.
The last line fixes the pilot's worst problem. Every new draft marks all older drafts for the case as superseded, and DocGen refuses to approve a superseded draft. At approval time, DocGen also asks the calculator whether its result version is still current. A draft with old figures cannot be approved, whichever one the specialist happens to open.
Events need the same care. Meridian's event platform delivers at least once, so the summary worker will occasionally see the same hardship.case.opened event twice. It records each event_id it has processed and ignores repeats.
Compensation for multi-step work
Some operations touch several systems. Creating a letter draft does three things: it creates the draft in DocGen, adds a note to the Atlas case, and writes an audit record. There is no transaction across three systems. If the second step fails after the first succeeded, the case has a draft that its notes do not mention.
The answer is compensation: for each step, define how to undo it or complete it later.
- Create the draft — if this fails, stop; nothing needs undoing.
- Record the audit entry locally, with an outbox message — the audit record and the message to send are written in one local transaction, so neither can exist without the other.
- Add the Atlas case note from the outbox — a relay retries until Atlas accepts it; the note is idempotent by draft ID.
- If the note still fails after 15 minutes, compensate — void the draft in DocGen and alert the service owner, so no draft exists without its case record.
The transactional outbox in step 2 is a well-known pattern: instead of calling another system directly inside a transaction, you write the message to a local table in the same transaction, and a separate process delivers it. It trades a small delay for a strong guarantee.
Partial results
Sometimes the right answer to a failure is a partial result. The rule is that a partial result must say it is partial. A summary that silently leaves out case notes because the notes API was down looks complete and misleads.
| Failure | Behaviour | What staff see |
|---|---|---|
| Notes API down during summary | Summarise records only | "Case notes unavailable at 10:42; this summary does not include notes." |
| Ledger slow on case open | Show precomputed narrative, retry facts | Facts table shows "Refreshing" with the last "as of" time |
| Reranker down | Switch to search mode | Top three policy passages, no generated answer |
| Calculator down | No letter draft | "Figures unavailable; letter drafting paused." Template mode is also blocked, because it needs figures |
| Check pass times out | Draft not shown | "Draft could not be verified. Try again or draft manually." |
The last row matters. An unverified draft is never displayed "just this once". The checks exist because the failures they catch are invisible to reviewers.
Part B of MER-07 records these semantics for every interface.
| Interface | Timeout | Retry | Idempotency | On failure | Compensation |
|---|---|---|---|---|---|
| Ledger facts | 1.5 s | Once, reads only | Natural | Show last "as of" | None needed |
| Atlas notes | 1.5 s | Once, reads only | Natural | Partial summary, labelled | None needed |
| Hardship Calculator | 1.5 s | Once | Natural | Pause drafting | None needed |
| Model draft | 15 s, deadline-aware | No, breaker instead | Not applicable | Template mode | None needed |
| DocGen create draft | 2 s | Yes, same key | Key from case, plan, figures, template | Error to staff | Supersede older drafts |
| Atlas case note | Outbox relay | Until 15 min | By draft ID | Alert | Void draft |
| Case opened event | Not applicable | Redelivered by platform | By event ID | Summary built on open | None needed |
Check your understanding
0 of 3 answered
1.Why does Meridian's draft idempotency key leave out the letter body?
2.The letter flow's step timeouts add up to 24.5 seconds against a 25-second requirement. Why pass a deadline down the chain as well?
3.The notes API is down when a summary is generated. What should the system do?