Course Content
AI System Design and Architecture
3 sections · 7 lessons
Mini Project: A Fault-Tolerant Inference Platform
You have been hired by an insurance claims processor. Their current system is one Flask file called app.py, 1,900 lines long, running on a single GPU box under someone's desk in Leeds. It classifies every uploaded claim document, extracts structured fields from the ones that are claim forms, and writes a plain-English summary for the claims that reach an adjuster.
Last Monday a partner switched on bulk uploads. Traffic went from 90 documents a minute to a peak of 4,000 a second. The box fell over at 09:12 and stayed down until 14:40 because restarting it reloads 19 GB of weights and the process kept being killed by the OOM reaper before it finished. Two hundred and eleven thousand documents were dropped with no record that they had ever arrived.
Your job is to rebuild it. Not to make it fancy — to make it survive Monday. Everything below is specified in numbers, because a design you cannot defend arithmetically is a guess.
The system you are building
| Capability | Model | Peak load | Per-item cost | Latency requirement |
|---|---|---|---|---|
classify | DistilBERT, 260 MB | 4,000 docs/s | CPU-servable, 3.1 ms at batch 16 | p99 < 250 ms, synchronous |
extract | LayoutLM-base, 1.4 GB | 120 docs/s | GPU, 42 ms at batch 8 | p99 < 900 ms, synchronous |
summarise | Llama-3-8B, 16 GB fp16 | 0.08 docs/s (about 5 a minute) | GPU, ~38 s per document | Asynchronous, 95% done in 5 min |
Constraints you must respect: a monthly infrastructure budget of 6,000 US dollars, an availability target of 99.9% for the synchronous endpoints, and the requirement that no accepted document is ever lost — not on a crash, not on a deploy, not on a zone failure.
Environment
You need Docker with Compose, a local Kubernetes (kind or k3d) or a small managed cluster, Python 3.11, a message broker (RabbitMQ is the easiest fit), Redis, Postgres, and a load generator — k6 or locust. Stub the models with functions that sleep for the measured times in the table above. Do not download real weights for this build. The architecture is the deliverable; a time.sleep(0.042) with the right latency distribution exercises every queue, timeout and breaker exactly as a real model would, and it lets you run the whole thing on a laptop.
Make the stub's latency realistic in one respect: give it a heavy tail. Return base × lognormal(0, 0.5) rather than a constant, so p99 sits roughly 3× above p50. A system tuned against constant latency will fall apart the first time it meets a real model.
Phase 1 — Decide the shape, and prove it
Write a one-page decision record. It must contain the following table, filled in with your own arithmetic, and a decision that follows from it.
| Capability | Peak load | Hardware class | Weights in memory | Replicas at 60% utilisation |
|---|---|---|---|---|
| classify | 4,000/s | CPU (8 vCPU) | 0.26 GB | ? |
| extract | 120/s | T4 GPU | 1.4 GB | ? |
| summarise | 0.08/s | A10G 24 GB | 16 GB | ? |
Here is the calculation for the first row so you can check your method. At batch 16, classify costs 3.1 ms per document, so one replica sustains 1000/3.1=322 docs/s. At 60% target utilisation each replica carries 193 docs/s, so you need ⌈4000/193⌉=21 replicas. On c6i.2xlarge at 0.34 US dollars an hour (AWS on-demand list price, as of September 2026) that is 7.14 an hour, or 5,212 a month — which already eats 87% of your budget, so the other two capabilities cannot possibly go on GPU instances alongside it.
That last sentence is the whole argument. A monolith puts summarise's 16 GB of weights inside every one of those 21 replicas, forcing all 21 onto 24 GB GPU instances at roughly 1.01 an hour: 21.13 an hour, or 15,425 a month — 2.6× over budget to serve a capability that receives one request every 12 seconds.
The split announces itself in the numbers: one capability's traffic is 50,000× another's, and one capability's weights are 60× another's. Those two facts, not architectural fashion, decide this.
Your decision record should land on three services plus a gateway, and should say explicitly what you are not splitting and why. Splitting preprocessing away from extract, for instance, would push a 588 KiB page tensor across a network boundary on every request — do not do it, and say so.
┌──────────────────────────┐ partner uploads ─►│ gateway (auth, quota, │ │ validation, routing) │ └───┬─────────┬────────┬───┘ sync ───────┘ sync ─┘ └─ async ▼ ▼ ▼ ┌────────────────┐ ┌───────────┐ ┌────────────────┐ │ classify-svc │ │extract-svc│ │ jobs.summarise │ │ 21 × CPU │ │ 2 × T4 │ │ (queue) │ └────────────────┘ └───────────┘ └───────┬────────┘ │ │ │ └───────┬───────┘ ┌───────┴────────┐ ▼ │ summarise │ ┌──────────────┐ │ workers 1..N │ │ Redis (L2) │ │ A10G, scale 1→6│ └──────────────┘ └───────┬────────┘ │ ▼ ┌───────┴──────────────────────────────────┐ │ Postgres: documents, jobs, results │ └──────────────────────────────────────────┘Phase 2 — The API contract
Build a versioned, resource-shaped API. The endpoints are not negotiable, because the shape is the point:
POST /v1/classifications sync → 200POST /v1/extractions sync → 200POST /v1/summaries async → 202 + job idGET /v1/summaries/{id} → status | resultDELETE /v1/summaries/{id} → cancelGET /v1/models → names, versions, limitsGET /healthz /readyz unversionedEvery successful response must carry id, model (name and version), usage (input tokens or pages, and compute milliseconds), and a warnings array. Every response — success or failure — must carry a request_id header. Every error must use a stable machine-readable code alongside a human message.
Two behaviours are graded hard. First, exceeding the documented input limit must return 413 or 422 with the actual limit in the message, never a silent truncation — silent truncation is how a model confidently scores the first paragraph of a 40-page claim. Second, POST /v1/summaries must honour an Idempotency-Key header, returning the original job id on a repeat rather than starting a second 38-second GPU job.
Prove the versioning works by making a genuinely breaking change on a branch: rename result.label to result.document_type. Show that /v1 clients still receive label while /v2 receives the new name, and write down which of your changes were additive (no bump needed) and which were breaking.
Phase 3 — Asynchrony and balancing
The summarise path is where Monday's outage lived: 38-second work behind a synchronous HTTP call with a 60-second load balancer timeout. Put a durable queue in front of it and return 202 immediately.
Requirements, each with a number you must justify:
- Durability. Persistent messages, publisher confirms, and a database row written before the enqueue. Demonstrate the guarantee: kill the broker mid-run and show that every accepted document still has a job row and eventually completes.
- Visibility timeout ≥ p99 processing time. If p99 is 74 seconds, a 60-second timeout means every slow job is redelivered while still running and processed twice. Set it to at least 120 s, or heartbeat-extend it.
- Prefetch = 1 on the summarise workers. With 38-second jobs, a worker that prefetches 10 messages holds six minutes of work hostage while its neighbours idle.
- Dead-letter queue with an alert. A corrupt PDF that fails deterministically must land in the DLQ after 3 attempts, not be requeued forever.
- Least-outstanding-requests balancing on the synchronous services, not round robin.
Demonstrate why the balancing rule matters with a measurement, not an assertion. Send a workload where 5% of documents take 20× as long as the rest, run it against round robin and against least-outstanding, and report both p99 figures. A 10× or better difference is the expected result.
Also compute the pooling benefit for your summarise fleet. With 4 workers at μ=1/38=0.0263 jobs/s each and λ=0.08 jobs/s, four private queues each at ρ=0.76 give an average wait of ρ/(μ−λi)=0.76/(0.0263−0.02)=120 s, while one shared queue with four servers gives about 21 s by the Erlang C formula — a 5.8× improvement from nothing but where the queue lives. State the numbers in your write-up.
Phase 4 — Autoscaling on a metric that means something
CPU utilisation is the wrong signal here: a summarise worker holding a GPU at 98% may show 9% CPU. Scale on queue depth per replica, derived from your SLA rather than guessed.
The target: 95% of summaries done within 5 minutes. Processing takes 38 s, so the queue budget is about 260 s. A worker completes 0.0263 jobs/s, so acceptable backlog per worker is 0.0263×260=6.8 — round to 6 jobs per replica.
Configure asymmetric behaviour and justify each number:
| Setting | Value | Why |
|---|---|---|
| Target | 6 queued jobs per replica | Derived above from the 5-minute SLA |
| Scale-up stabilisation | 30 s | React before the backlog compounds |
| Scale-up policy | +100% per 60 s | Recovery requires overshoot, not exact sizing |
| Scale-down stabilisation | 600 s | Cold start is expensive; retreat slowly |
| Scale-down policy | −1 replica per 120 s | Prevents thrashing |
| minReplicas / maxReplicas | 1 / 6 | 6 × 1.01/hr × 730 = 4,424 a month — the most a runaway autoscaler can spend here; check it against the 6,000 budget together with the other services |
Then measure your cold start honestly and write the recovery arithmetic. If evaluation takes 60 s, node provisioning 90 s, image pull 32 s and model load 45 s, that is 227 s before new capacity serves. During a jump from 0.08 to 0.4 jobs/s with 4 workers (capacity 0.105 jobs/s) the deficit is 0.295 jobs/s, so the backlog reaches 0.295×227=67 jobs before new capacity arrives. Merely keeping up with 0.4 jobs/s needs 16 workers (0.4×38=15.2), and draining the backlog needs more than that — far beyond the maxReplicas of 6 above. So decide in writing what happens in a surge like this: raise the cap and pay for it, or let summaries miss the 5-minute target while the queue absorbs the burst. Show the graph.
Phase 5 — Caching, keyed correctly
Two tiers: an in-process LRU for the hottest keys and Redis shared across replicas.
The cache key must be sha256(normalised_input) + ":" + model_name + "@" + model_version. Model version in the key is the single most important line in this phase. It means deploying v8 makes every v7 entry unreachable without a flush, and a rollback to v7 finds its cache still warm. A flush-on-deploy design sends 100% of traffic to cold replicas at the exact moment you are also rolling pods — a routine release becomes an outage.
Requirements: a Redis timeout of 50 ms with the client degrading to direct computation on any error; single-flight locking so that when a hot key expires, one request recomputes and the rest wait rather than 300 concurrent misses hitting the model; and a reported hit rate per tier.
Then produce the effect calculation for your measured hit rate. At 52% overall on a 45 ms model with a 2 ms cache, effective latency is 0.52×2+0.48×45=22.6 ms and model load falls by 52%, taking the classify fleet from 21 replicas to 11 and saving roughly 2,480 a month. Show your own equivalent numbers.
Finally, demonstrate the warm-up problem. Deploy with a cold cache under full load and record what happens; then implement a warm-up that replays the top 2,000 keys from the last 24 hours before a replica reports ready, and record the difference. A fleet sized for a 52% hit rate faces 1/0.48=2.08× its provisioned load when cold.
Phase 6 — Make it survive
Implement all six, and test all six:
| Mechanism | Concrete requirement | Test you must run |
|---|---|---|
| Explicit timeouts | Every outbound call, set from measured p99 | Blackhole Redis; service must degrade, not error |
| Retry with full jitter | One layer only, max 3, deadline-aware | Show total downstream calls per client request never exceeds 3 |
| Circuit breaker | Opens at 50% failures over 20 calls, half-open after 30 s | Kill extract-svc; gateway must fail in ~1 ms, not 6 s |
| Bulkheads | Separate pools per dependency | Hang the DB; classify must keep serving |
| Graceful degradation | Documented fallback per dependency | Kill the summariser; classify and extract unaffected |
| Health probes | Liveness with zero dependencies; readiness gated on a warm-up inference; drain on SIGTERM | Rolling restart must produce zero 5xx |
Size your replica count for the degraded state, not the happy path. For classify at 4,000 docs/s with 193 docs/s per replica at 60% utilisation, ask what happens after losing an availability zone. If your 21 replicas are spread across 3 zones, losing one leaves 14 carrying 286 docs/s each — 89% utilisation, where the ρ/(1−ρ) queue factor is 7.9 against 1.5 at 60%, so waits are about 5× worse. To hold 75% after a zone loss you need ⌈4000/(322×0.75)⌉=17 surviving replicas, hence 26 total across 3 zones. That difference — 21 versus 26 — is what your 99.9% target actually costs.
Phase 7 — Shipping a new model
Roll out classify@v8 using a canary, and justify it with the blast-radius arithmetic. Suppose v8 has an 8% error rate and takes 10 minutes to detect. At 100% traffic that is 10×1.00×0.08=0.8 error-minutes; at a 5% canary it is 10×0.05×0.08=0.04 error-minutes, twenty times less. Against a 99.9% monthly SLO the error budget is 43.2 minutes, so the canary lets you ship far more often before the budget is gone.
Your promotion gate must check three things at each step (5% → 25% → 50% → 100%), and must automatically roll back on any breach:
- Error rate of the canary within 1.5× the incumbent's.
- p99 latency of the canary within 1.2× the incumbent's.
- Prediction distribution: mean confidence and per-class rates within a stated tolerance of the incumbent, measured on the same traffic.
The third gate is the one that distinguishes a model rollout from a software rollout. A model that returns 200 OK for everything while its mean confidence falls from 0.91 to 0.68 is broken and will pass both of the first two checks.
Deliverables and how they are judged
claims-platform/ docs/adr-001-service-boundaries.md ← Phase 1, with the numbers docs/latency-budget.md ← per-stage p50/p99 table docs/runbook.md ← what to do when each thing breaks gateway/ auth, quota, routing, breakers services/classify/ api + batching + two-tier cache services/extract/ api + batching workers/summarise/ consumer, idempotent, heartbeating deploy/ k8s manifests, HPA, canary weights load/ k6 scripts: steady, burst, heavy-tail, zone-loss chaos/ scripts that kill redis / db / a zone| Area | Points | What earns full marks |
|---|---|---|
| Architecture decision | 20 | A decision record whose conclusion follows from a capacity and cost table, including what you chose not to split |
| API design | 15 | Resource-shaped, versioned independently of the model, limits enforced with 4xx, idempotency honoured |
| Async and balancing | 20 | No document lost under broker restart; measured p99 difference between balancing strategies |
| Autoscaling | 10 | Target derived from the SLA; asymmetric windows; cold-start recovery measured |
| Caching | 15 | Version in the key, single-flight, degrades on Redis failure, warm-up demonstrated |
| Fault tolerance | 15 | All six mechanisms present and each proven by a chaos test |
| Deployment | 5 | Canary with an automatic quality gate, not just an error-rate gate |
Where this build usually goes wrong
Four failure modes account for most of the marks lost, and all four are avoidable.
Building the microservices first and the numbers afterwards. If your decision record was written after the code, it will not survive a question like "why 21 replicas and not 12?". Write the capacity table before you open an editor; it takes twenty minutes and it determines everything else.
Write the capacity table before you open an editor. Every other decision in this build is downstream of it.
Testing only the steady state. A system that handles 4,000 docs/s and a system that can get from 400 to 4,000 docs/s without a seven-minute breach are different systems. Your load scripts must include a step change, and your report must include the recovery time.
Fault tolerance that is never exercised. A circuit breaker you have not tripped is a configuration file, not a safety mechanism. Every one of the six mechanisms needs a chaos script that proves it, and running those scripts is what turns beliefs into evidence.
Optimising the model instead of the queue. When p99 disappoints, the instinct is to make inference faster. Break the budget down by stage first: on a fleet at 89% utilisation, queue wait will be the largest term, and two extra replicas will beat a fortnight of optimisation work.
When you are finished, the test is not whether the demo works. It is whether you can be handed the Monday scenario — 90 documents a minute becoming 4,000 a second without warning — and say precisely what your system does in the first 30 seconds, the first 5 minutes, and the first hour, with a number attached to each answer. That is what separates a system that has been built from a system that has been designed.