Enterprise AI Solutions Architecture

Course Content

Enterprise AI Solutions Architecture

13 sections · 29 lessons

Failure Semantics: Idempotency, Timeouts and Compensation


In the second week of the pilot, a hardship specialist clicked "Draft letter", waited, and clicked again. Meanwhile the gateway timed out the first request and retried it. DocGen received three requests and created three drafts for the same case. One of them had been generated before the specialist changed the plan length from six months to nine, so it carried the old figures. It was the one she opened and approved. QA caught it in the 5% sample, two days after it was sent.

None of this was about AI. It was a classic distributed systems failure: no timeouts that fitted together, no idempotency and no rule about which draft was current. But AI makes these failures more likely, because model calls are slow. A call that takes 12 seconds invites double-clicks, hits timeouts and triggers retries far more often than a call that takes 50 milliseconds.

This lesson defines Meridian's failure semantics: what every interface does when it is slow, down or called twice. It produces part B of MER-07.

How the three-drafts bug is closedDouble-clickplus agateway retrySame key:case, plan,figures, templateDocGen returns theone existing draftPlan change:new key, oldsupersededApproval recheckscalculator versionThe letter body is left out of the key; the model words it differently every time.
Slow model calls invite double-clicks and retries, and three simple, testable controls stop them becoming wrong letters.

Timeouts that add up

A timeout is not the same as a p95. The p95 describes normal behaviour; a timeout says when to give up. Set timeouts from two directions: long enough that normal slow calls finish, and short enough that the sum of sequential steps fits the user-facing deadline.

The letter draft has a 25-second p95 requirement. Here is how its budget is divided.

StepRunsp95Timeout
Calculator, account facts, notesIn parallel450 ms1.5 s
Model draft, about 900 tokensSequential11 s15 s
Check pass: figures, paragraphs, claimsSequential4 s6 s
DocGen create draftSequential300 ms2 s
Worst case before giving up24.5 s

The worst case, with every step at its timeout, still fits under 25 seconds. The design also passes a deadline down the chain: each step receives the time remaining and uses the smaller of that and its own timeout. If the model takes 14 seconds, the check pass gets only what is left, and the request fails cleanly rather than running over.

Idempotency

An operation is idempotent if doing it twice has the same effect as doing it once. Reads are naturally idempotent. Draft creation is not, unless you make it so. The standard technique is an idempotency key: a value derived from what makes the request unique, sent with the request, so the receiver can recognise a repeat.

Python
import hashlibimport jsondef draft_key(case_id: str, plan: dict, calc_version: str, template_id: str) -> str:    """Same case, plan, figures and template means the same draft."""    material = json.dumps(        {"case": case_id, "plan": plan, "calc": calc_version, "tpl": template_id},        sort_keys=True,    )    return hashlib.sha256(material.encode()).hexdigest()[:32]def create_draft_once(docgen, ctx, plan: dict, calc: dict, template_id: str, body: str) -> str:    key = draft_key(ctx.case_id, plan, calc["version"], template_id)    draft_id = docgen.create_draft(        case_id=ctx.case_id, template_id=template_id, body=body,        figures=calc["figures"], idempotency_key=key, token=ctx.obo_token,    )    docgen.supersede_others(case_id=ctx.case_id, keep=draft_id, token=ctx.obo_token)    return draft_id

The key is built from the case, the plan terms, the calculator result version and the template. It deliberately leaves out the letter body, because the model writes different words each time; a retry must not count as a new draft just because the wording changed. DocGen stores the key and, if it sees the same key again, returns the existing draft instead of creating another. That server-side check is the real guarantee; a check only in the caller would still race when two requests arrive together.

The last line fixes the pilot's worst problem. Every new draft marks all older drafts for the case as superseded, and DocGen refuses to approve a superseded draft. At approval time, DocGen also asks the calculator whether its result version is still current. A draft with old figures cannot be approved, whichever one the specialist happens to open.

Events need the same care. Meridian's event platform delivers at least once, so the summary worker will occasionally see the same hardship.case.opened event twice. It records each event_id it has processed and ignores repeats.

Compensation for multi-step work

Some operations touch several systems. Creating a letter draft does three things: it creates the draft in DocGen, adds a note to the Atlas case, and writes an audit record. There is no transaction across three systems. If the second step fails after the first succeeded, the case has a draft that its notes do not mention.

The answer is compensation: for each step, define how to undo it or complete it later.

  1. Create the draft — if this fails, stop; nothing needs undoing.
  2. Record the audit entry locally, with an outbox message — the audit record and the message to send are written in one local transaction, so neither can exist without the other.
  3. Add the Atlas case note from the outbox — a relay retries until Atlas accepts it; the note is idempotent by draft ID.
  4. If the note still fails after 15 minutes, compensate — void the draft in DocGen and alert the service owner, so no draft exists without its case record.

The transactional outbox in step 2 is a well-known pattern: instead of calling another system directly inside a transaction, you write the message to a local table in the same transaction, and a separate process delivers it. It trades a small delay for a strong guarantee.

Partial results

Sometimes the right answer to a failure is a partial result. The rule is that a partial result must say it is partial. A summary that silently leaves out case notes because the notes API was down looks complete and misleads.

FailureBehaviourWhat staff see
Notes API down during summarySummarise records only"Case notes unavailable at 10:42; this summary does not include notes."
Ledger slow on case openShow precomputed narrative, retry factsFacts table shows "Refreshing" with the last "as of" time
Reranker downSwitch to search modeTop three policy passages, no generated answer
Calculator downNo letter draft"Figures unavailable; letter drafting paused." Template mode is also blocked, because it needs figures
Check pass times outDraft not shown"Draft could not be verified. Try again or draft manually."

The last row matters. An unverified draft is never displayed "just this once". The checks exist because the failures they catch are invisible to reviewers.

Part B of MER-07 records these semantics for every interface.

InterfaceTimeoutRetryIdempotencyOn failureCompensation
Ledger facts1.5 sOnce, reads onlyNaturalShow last "as of"None needed
Atlas notes1.5 sOnce, reads onlyNaturalPartial summary, labelledNone needed
Hardship Calculator1.5 sOnceNaturalPause draftingNone needed
Model draft15 s, deadline-awareNo, breaker insteadNot applicableTemplate modeNone needed
DocGen create draft2 sYes, same keyKey from case, plan, figures, templateError to staffSupersede older drafts
Atlas case noteOutbox relayUntil 15 minBy draft IDAlertVoid draft
Case opened eventNot applicableRedelivered by platformBy event IDSummary built on openNone needed

Check your understanding

0 of 3 answered

1.Why does Meridian's draft idempotency key leave out the letter body?

2.The letter flow's step timeouts add up to 24.5 seconds against a 25-second requirement. Why pass a deadline down the chain as well?

3.The notes API is down when a summary is generated. What should the system do?