AI System Design and Architecture

Mini Project: A Fault-Tolerant Inference Platform


You have been hired by an insurance claims processor. Their current system is one Flask file called app.py, 1,900 lines long, running on a single GPU box under someone's desk in Leeds. It classifies every uploaded claim document, extracts structured fields from the ones that are claim forms, and writes a plain-English summary for the claims that reach an adjuster.

Last Monday a partner switched on bulk uploads. Traffic went from 90 documents a minute to a peak of 4,000 a second. The box fell over at 09:12 and stayed down until 14:40 because restarting it reloads 19 GB of weights and the process kept being killed by the OOM reaper before it finished. Two hundred and eleven thousand documents were dropped with no record that they had ever arrived.

Your job is to rebuild it. Not to make it fancy — to make it survive Monday. Everything below is specified in numbers, because a design you cannot defend arithmetically is a guess.

What has to be true before you delete app.pyOne claimdocument,three modelsAsync job id,not a 38 s blockAutoscale on queue depthCache keyed on document hashTimeout andbreaker per modelCanary a newmodel, then widen
Splitting the 1,900-line file is the easy half; the platform is the timeouts, the queue and the way back out.

The system you are building

CapabilityModelPeak loadPer-item costLatency requirement
classifyDistilBERT, 260 MB4,000 docs/sCPU-servable, 3.1 ms at batch 16p99 < 250 ms, synchronous
extractLayoutLM-base, 1.4 GB120 docs/sGPU, 42 ms at batch 8p99 < 900 ms, synchronous
summariseLlama-3-8B, 16 GB fp160.08 docs/s (about 5 a minute)GPU, ~38 s per documentAsynchronous, 95% done in 5 min

Constraints you must respect: a monthly infrastructure budget of 6,000 US dollars, an availability target of 99.9% for the synchronous endpoints, and the requirement that no accepted document is ever lost — not on a crash, not on a deploy, not on a zone failure.

Environment

You need Docker with Compose, a local Kubernetes (kind or k3d) or a small managed cluster, Python 3.11, a message broker (RabbitMQ is the easiest fit), Redis, Postgres, and a load generator — k6 or locust. Stub the models with functions that sleep for the measured times in the table above. Do not download real weights for this build. The architecture is the deliverable; a time.sleep(0.042) with the right latency distribution exercises every queue, timeout and breaker exactly as a real model would, and it lets you run the whole thing on a laptop.

Make the stub's latency realistic in one respect: give it a heavy tail. Return base × lognormal(0, 0.5) rather than a constant, so p99 sits roughly 3× above p50. A system tuned against constant latency will fall apart the first time it meets a real model.

Phase 1 — Decide the shape, and prove it

Write a one-page decision record. It must contain the following table, filled in with your own arithmetic, and a decision that follows from it.

CapabilityPeak loadHardware classWeights in memoryReplicas at 60% utilisation
classify4,000/sCPU (8 vCPU)0.26 GB?
extract120/sT4 GPU1.4 GB?
summarise0.08/sA10G 24 GB16 GB?

Here is the calculation for the first row so you can check your method. At batch 16, classify costs 3.1 ms per document, so one replica sustains 1000/3.1=3221000/3.1 = 322 docs/s. At 60% target utilisation each replica carries 193 docs/s, so you need ⌈4000/193⌉=21\lceil 4000/193 \rceil = 21 replicas. On c6i.2xlarge at 0.34 US dollars an hour (AWS on-demand list price, as of September 2026) that is 7.14 an hour, or 5,212 a month — which already eats 87% of your budget, so the other two capabilities cannot possibly go on GPU instances alongside it.

That last sentence is the whole argument. A monolith puts summarise's 16 GB of weights inside every one of those 21 replicas, forcing all 21 onto 24 GB GPU instances at roughly 1.01 an hour: 21.13 an hour, or 15,425 a month — 2.6× over budget to serve a capability that receives one request every 12 seconds.

The split announces itself in the numbers: one capability's traffic is 50,000× another's, and one capability's weights are 60× another's. Those two facts, not architectural fashion, decide this.

Your decision record should land on three services plus a gateway, and should say explicitly what you are not splitting and why. Splitting preprocessing away from extract, for instance, would push a 588 KiB page tensor across a network boundary on every request — do not do it, and say so.

Text
                     ┌──────────────────────────┐   partner uploads ─►│ gateway  (auth, quota,   │                     │ validation, routing)     │                     └───┬─────────┬────────┬───┘             sync ───────┘   sync ─┘        └─ async                 ▼               ▼                ▼        ┌────────────────┐ ┌───────────┐   ┌────────────────┐        │ classify-svc   │ │extract-svc│   │ jobs.summarise │        │ 21 × CPU       │ │ 2 × T4    │   │  (queue)       │        └────────────────┘ └───────────┘   └───────┬────────┘                 │               │                 │                 └───────┬───────┘         ┌───────┴────────┐                         ▼                 │ summarise      │                 ┌──────────────┐          │ workers 1..N   │                 │ Redis (L2)   │          │ A10G, scale 1→6│                 └──────────────┘          └───────┬────────┘                         │                         ▼                 ┌───────┴──────────────────────────────────┐                 │ Postgres: documents, jobs, results       │                 └──────────────────────────────────────────┘

Phase 2 — The API contract

Build a versioned, resource-shaped API. The endpoints are not negotiable, because the shape is the point:

Text
POST   /v1/classifications          sync   → 200POST   /v1/extractions              sync   → 200POST   /v1/summaries                async  → 202 + job idGET    /v1/summaries/{id}                  → status | resultDELETE /v1/summaries/{id}                  → cancelGET    /v1/models                          → names, versions, limitsGET    /healthz  /readyz            unversioned

Every successful response must carry id, model (name and version), usage (input tokens or pages, and compute milliseconds), and a warnings array. Every response — success or failure — must carry a request_id header. Every error must use a stable machine-readable code alongside a human message.

Two behaviours are graded hard. First, exceeding the documented input limit must return 413 or 422 with the actual limit in the message, never a silent truncation — silent truncation is how a model confidently scores the first paragraph of a 40-page claim. Second, POST /v1/summaries must honour an Idempotency-Key header, returning the original job id on a repeat rather than starting a second 38-second GPU job.

Prove the versioning works by making a genuinely breaking change on a branch: rename result.label to result.document_type. Show that /v1 clients still receive label while /v2 receives the new name, and write down which of your changes were additive (no bump needed) and which were breaking.

Phase 3 — Asynchrony and balancing

The summarise path is where Monday's outage lived: 38-second work behind a synchronous HTTP call with a 60-second load balancer timeout. Put a durable queue in front of it and return 202 immediately.

Requirements, each with a number you must justify:

  • Durability. Persistent messages, publisher confirms, and a database row written before the enqueue. Demonstrate the guarantee: kill the broker mid-run and show that every accepted document still has a job row and eventually completes.
  • Visibility timeout ≥ p99 processing time. If p99 is 74 seconds, a 60-second timeout means every slow job is redelivered while still running and processed twice. Set it to at least 120 s, or heartbeat-extend it.
  • Prefetch = 1 on the summarise workers. With 38-second jobs, a worker that prefetches 10 messages holds six minutes of work hostage while its neighbours idle.
  • Dead-letter queue with an alert. A corrupt PDF that fails deterministically must land in the DLQ after 3 attempts, not be requeued forever.
  • Least-outstanding-requests balancing on the synchronous services, not round robin.

Demonstrate why the balancing rule matters with a measurement, not an assertion. Send a workload where 5% of documents take 20× as long as the rest, run it against round robin and against least-outstanding, and report both p99 figures. A 10× or better difference is the expected result.

Also compute the pooling benefit for your summarise fleet. With 4 workers at μ=1/38=0.0263\mu = 1/38 = 0.0263 jobs/s each and λ=0.08\lambda = 0.08 jobs/s, four private queues each at ρ=0.76\rho = 0.76 give an average wait of ρ/(μ−λi)=0.76/(0.0263−0.02)=120\rho/(\mu-\lambda_i) = 0.76/(0.0263-0.02) = 120 s, while one shared queue with four servers gives about 21 s by the Erlang C formula — a 5.8× improvement from nothing but where the queue lives. State the numbers in your write-up.

Phase 4 — Autoscaling on a metric that means something

CPU utilisation is the wrong signal here: a summarise worker holding a GPU at 98% may show 9% CPU. Scale on queue depth per replica, derived from your SLA rather than guessed.

The target: 95% of summaries done within 5 minutes. Processing takes 38 s, so the queue budget is about 260 s. A worker completes 0.0263 jobs/s, so acceptable backlog per worker is 0.0263×260=6.80.0263 \times 260 = 6.8 — round to 6 jobs per replica.

Configure asymmetric behaviour and justify each number:

SettingValueWhy
Target6 queued jobs per replicaDerived above from the 5-minute SLA
Scale-up stabilisation30 sReact before the backlog compounds
Scale-up policy+100% per 60 sRecovery requires overshoot, not exact sizing
Scale-down stabilisation600 sCold start is expensive; retreat slowly
Scale-down policy−1 replica per 120 sPrevents thrashing
minReplicas / maxReplicas1 / 66 × 1.01/hr × 730 = 4,424 a month — the most a runaway autoscaler can spend here; check it against the 6,000 budget together with the other services

Then measure your cold start honestly and write the recovery arithmetic. If evaluation takes 60 s, node provisioning 90 s, image pull 32 s and model load 45 s, that is 227 s before new capacity serves. During a jump from 0.08 to 0.4 jobs/s with 4 workers (capacity 0.105 jobs/s) the deficit is 0.295 jobs/s, so the backlog reaches 0.295×227=670.295 \times 227 = 67 jobs before new capacity arrives. Merely keeping up with 0.4 jobs/s needs 16 workers (0.4×38=15.20.4 \times 38 = 15.2), and draining the backlog needs more than that — far beyond the maxReplicas of 6 above. So decide in writing what happens in a surge like this: raise the cap and pay for it, or let summaries miss the 5-minute target while the queue absorbs the burst. Show the graph.

Phase 5 — Caching, keyed correctly

Two tiers: an in-process LRU for the hottest keys and Redis shared across replicas.

The cache key must be sha256(normalised_input) + ":" + model_name + "@" + model_version. Model version in the key is the single most important line in this phase. It means deploying v8 makes every v7 entry unreachable without a flush, and a rollback to v7 finds its cache still warm. A flush-on-deploy design sends 100% of traffic to cold replicas at the exact moment you are also rolling pods — a routine release becomes an outage.

Requirements: a Redis timeout of 50 ms with the client degrading to direct computation on any error; single-flight locking so that when a hot key expires, one request recomputes and the rest wait rather than 300 concurrent misses hitting the model; and a reported hit rate per tier.

Then produce the effect calculation for your measured hit rate. At 52% overall on a 45 ms model with a 2 ms cache, effective latency is 0.52×2+0.48×45=22.60.52 \times 2 + 0.48 \times 45 = 22.6 ms and model load falls by 52%, taking the classify fleet from 21 replicas to 11 and saving roughly 2,480 a month. Show your own equivalent numbers.

Finally, demonstrate the warm-up problem. Deploy with a cold cache under full load and record what happens; then implement a warm-up that replays the top 2,000 keys from the last 24 hours before a replica reports ready, and record the difference. A fleet sized for a 52% hit rate faces 1/0.48=2.08×1/0.48 = 2.08\times its provisioned load when cold.

Phase 6 — Make it survive

Implement all six, and test all six:

MechanismConcrete requirementTest you must run
Explicit timeoutsEvery outbound call, set from measured p99Blackhole Redis; service must degrade, not error
Retry with full jitterOne layer only, max 3, deadline-awareShow total downstream calls per client request never exceeds 3
Circuit breakerOpens at 50% failures over 20 calls, half-open after 30 sKill extract-svc; gateway must fail in ~1 ms, not 6 s
BulkheadsSeparate pools per dependencyHang the DB; classify must keep serving
Graceful degradationDocumented fallback per dependencyKill the summariser; classify and extract unaffected
Health probesLiveness with zero dependencies; readiness gated on a warm-up inference; drain on SIGTERMRolling restart must produce zero 5xx

Size your replica count for the degraded state, not the happy path. For classify at 4,000 docs/s with 193 docs/s per replica at 60% utilisation, ask what happens after losing an availability zone. If your 21 replicas are spread across 3 zones, losing one leaves 14 carrying 286 docs/s each — 89% utilisation, where the ρ/(1−ρ)\rho/(1-\rho) queue factor is 7.9 against 1.5 at 60%, so waits are about 5× worse. To hold 75% after a zone loss you need ⌈4000/(322×0.75)⌉=17\lceil 4000/(322 \times 0.75) \rceil = 17 surviving replicas, hence 26 total across 3 zones. That difference — 21 versus 26 — is what your 99.9% target actually costs.

Phase 7 — Shipping a new model

Roll out classify@v8 using a canary, and justify it with the blast-radius arithmetic. Suppose v8 has an 8% error rate and takes 10 minutes to detect. At 100% traffic that is 10×1.00×0.08=0.810 \times 1.00 \times 0.08 = 0.8 error-minutes; at a 5% canary it is 10×0.05×0.08=0.0410 \times 0.05 \times 0.08 = 0.04 error-minutes, twenty times less. Against a 99.9% monthly SLO the error budget is 43.2 minutes, so the canary lets you ship far more often before the budget is gone.

Your promotion gate must check three things at each step (5% → 25% → 50% → 100%), and must automatically roll back on any breach:

  • Error rate of the canary within 1.5× the incumbent's.
  • p99 latency of the canary within 1.2× the incumbent's.
  • Prediction distribution: mean confidence and per-class rates within a stated tolerance of the incumbent, measured on the same traffic.

The third gate is the one that distinguishes a model rollout from a software rollout. A model that returns 200 OK for everything while its mean confidence falls from 0.91 to 0.68 is broken and will pass both of the first two checks.

Deliverables and how they are judged

Text
claims-platform/  docs/adr-001-service-boundaries.md      ← Phase 1, with the numbers  docs/latency-budget.md                  ← per-stage p50/p99 table  docs/runbook.md                         ← what to do when each thing breaks  gateway/            auth, quota, routing, breakers  services/classify/  api + batching + two-tier cache  services/extract/   api + batching  workers/summarise/  consumer, idempotent, heartbeating  deploy/             k8s manifests, HPA, canary weights  load/               k6 scripts: steady, burst, heavy-tail, zone-loss  chaos/              scripts that kill redis / db / a zone
AreaPointsWhat earns full marks
Architecture decision20A decision record whose conclusion follows from a capacity and cost table, including what you chose not to split
API design15Resource-shaped, versioned independently of the model, limits enforced with 4xx, idempotency honoured
Async and balancing20No document lost under broker restart; measured p99 difference between balancing strategies
Autoscaling10Target derived from the SLA; asymmetric windows; cold-start recovery measured
Caching15Version in the key, single-flight, degrades on Redis failure, warm-up demonstrated
Fault tolerance15All six mechanisms present and each proven by a chaos test
Deployment5Canary with an automatic quality gate, not just an error-rate gate

Where this build usually goes wrong

Four failure modes account for most of the marks lost, and all four are avoidable.

Building the microservices first and the numbers afterwards. If your decision record was written after the code, it will not survive a question like "why 21 replicas and not 12?". Write the capacity table before you open an editor; it takes twenty minutes and it determines everything else.

Write the capacity table before you open an editor. Every other decision in this build is downstream of it.

Testing only the steady state. A system that handles 4,000 docs/s and a system that can get from 400 to 4,000 docs/s without a seven-minute breach are different systems. Your load scripts must include a step change, and your report must include the recovery time.

Fault tolerance that is never exercised. A circuit breaker you have not tripped is a configuration file, not a safety mechanism. Every one of the six mechanisms needs a chaos script that proves it, and running those scripts is what turns beliefs into evidence.

Optimising the model instead of the queue. When p99 disappoints, the instinct is to make inference faster. Break the budget down by stage first: on a fleet at 89% utilisation, queue wait will be the largest term, and two extra replicas will beat a fortnight of optimisation work.

When you are finished, the test is not whether the demo works. It is whether you can be handed the Monday scenario — 90 documents a minute becoming 4,000 a second without warning — and say precisely what your system does in the first 30 seconds, the first 5 minutes, and the first hour, with a number attached to each answer. That is what separates a system that has been built from a system that has been designed.