AI System Design and Architecture

Monolithic vs. Microservice Architectures


A four-person team ships a content moderation service. It is one Python process. Inside it live three models: a DistilBERT text-toxicity classifier, a CLIP image classifier for unsafe pictures, and a Llama-3-8B model that writes human-readable explanations when a user appeals a takedown. All three are loaded at process start. The service works, it is easy to reason about, and it took six weeks.

Then the product launches in a second country and text traffic goes from 200 requests per second to 3,000. Image traffic barely moves. Appeals stay at roughly one every two seconds. The team does the only thing the architecture allows: they add replicas of the whole process. Because Llama-3-8B in fp16 needs 16 GB of GPU memory, every replica must run on a GPU instance with at least 24 GB. They end up running nine copies of an 8-billion-parameter model to serve 0.5 requests per second, because that model happens to share a process with the one that is actually busy.

That is the moment architecture stops being an abstract word. The system is not slow and it is not broken. It is shaped wrong: the unit you can scale is not the unit that needs scaling.

Three models in one process, or three servicesOne process, three models• One image, one deploy, one log stream• All three share the GPU and the RAM• Llama-3-8B forcesthe whole box to be big• Scaling the appeal path scales OCR tooThree services, three lifecycles• Each model scales on its own traffic• CPU boxes for text, GPU only for Llama• A network hop and aretry policy per call• Three pipelines andthree things to page on
Splitting buys independent scaling and charges you in latency, deploys and surface area — which is why modular monoliths ship.

What architecture actually decides

System architecture is the set of decisions about what the separately deployable pieces are, what talks to what, and where state lives. Everything else — which web framework, which ORM, which logging library — you can change on a Tuesday afternoon. Architecture you cannot, because changing it means changing every deployment, every on-call runbook, and often the team structure.

AI systems put unusual pressure on those decisions, for reasons that ordinary CRUD web applications never face:

  • The artefact is enormous. A container for a web API is 200 MB. A container carrying model weights is 2–20 GB. Deployment time, image pull time, and cold-start time are all dominated by weights, not code.
  • Compute is heterogeneous. Some models want a GPU, some run happily on four vCPUs, some want a lot of RAM and no GPU at all. Packing them into one process forces every replica onto the most expensive hardware any one of them needs.
  • Capacity is measured in model-seconds, not queries. A single request may occupy a GPU for 40 ms exclusively. You cannot serve your way out of that with more threads.
  • The model changes independently of the code. Retraining happens weekly; the API contract may not change for a year. If the two are welded together, every retrain is a full application deploy.

Architecture is the answer to one question: what is the smallest thing you can deploy, scale, and break on its own?

The monolith

A monolithic AI service is a single deployable unit that contains the API layer, the preprocessing, every model, and the postprocessing. Components talk by function call inside one process.

Text
                 ┌─────────────────────────────────────┐   client ─────► │  moderation-service  (one process)  │                 │                                     │                 │   HTTP layer                        │                 │      │                              │                 │      ├─► preprocess()               │                 │      ├─► text_model.predict()       │                 │      ├─► image_model.predict()      │                 │      ├─► llm.generate()             │                 │      └─► postprocess()              │                 │                                     │                 │   loaded weights: 17.9 GB GPU       │                 └─────────────────────────────────────┘                                  │                            ┌─────▼─────┐                            │ Postgres  │                            └───────────┘

What it gets right

The advantages are not sentimental, they are measurable. Calling text_model.predict() from the request handler costs roughly 50 microseconds of overhead. The equivalent call to another service over HTTP inside the same VPC costs about 1.2 ms at p50 and 9 ms at p99 once you count connection handling, serialisation and the remote server's own queue. That is a 24× to 180× difference in overhead, per hop.

Passing data is free. A preprocessed image tensor of shape 224×224×3 in float32 is 150,528 floats, or 602,112 bytes — 588 KiB. Inside a process you pass a pointer to it. Across a service boundary you must serialise it. If you naively JSON-encode it as base64, it inflates by 4/3 to 802,816 bytes and costs several milliseconds of CPU on both ends. Teams discover this the hard way.

You also get one deploy, one log stream, one place to set a breakpoint, and transactional consistency for free because there is one database connection and one process.

What it gets wrong

ProblemConcrete symptom
Coupled scalingNine replicas of a 16 GB LLM to serve 0.5 req/s, because it shares a process with the busy text model.
Hardware lowest-common-denominatorEvery replica needs a 24 GB GPU even though 95% of requests only touch a CPU-servable classifier.
Blast radiusAn out-of-memory error while decoding one malformed PNG kills the process, taking text moderation and appeals down with it.
Deploy frictionChanging a regex in the text path means rebuilding and re-rolling an 18 GB image across nine replicas, each with a 90 s model warm-up.
Dependency conflictsOne model needs transformers==4.36, the vision stack needs torch==2.1 compiled against a CUDA version the other cannot use. One environment, one winner.

Microservices

Split by capability. Each model, or each closely related group of models, becomes its own deployable service with its own image, its own scaling policy, and its own hardware.

Text
                     ┌──────────────┐   client ──────────►│  API gateway │                     └──┬────┬───┬──┘             ┌──────────┘    │   └──────────────┐             ▼               ▼                  ▼  ┌────────────────┐ ┌───────────────┐ ┌──────────────────┐  │ text-toxicity  │ │  image-nsfw   │ │ appeal-explainer │  │ 12 × c6i.2xl   │ │ 2 × g4dn.xl   │ │ 2 × g5.xlarge    │  │ CPU, 0.3 GB    │ │ T4, 1.6 GB    │ │ A10G, 16 GB      │  │ 3,000 req/s    │ │ 200 img/s     │ │ 0.5 req/s        │  └────────────────┘ └───────────────┘ └──────────────────┘             │               │                  │             └───────────────┴──────────────────┘                             ▼                    ┌────────────────┐                    │  shared store  │                    └────────────────┘

The cost arithmetic, done properly

Take the same traffic: 3,000 text req/s, 200 img/s, 0.5 appeal req/s at peak. Measured throughput per replica, at 60% target utilisation so there is headroom for bursts. Prices are AWS us-east-1 on-demand list prices as of September 2026; check the current rate card before you reuse them.

WorkloadInstanceUSD/hr (on-demand)Safe throughput per replicaReplicas needed
Monolith (all three models)g5.xlarge, A10G 24 GB1.006350 text req/s⌈3000/350⌉ = 9
text-toxicity onlyc6i.2xlarge, 8 vCPU0.340250 req/s⌈3000/250⌉ = 12
image-nsfw onlyg4dn.xlarge, T4 16 GB0.526130 img/s⌈200/130⌉ = 2
appeal-explainer onlyg5.xlarge, A10G 24 GB1.0061.5 req/s2 (for redundancy, not load)

Monolith: 9 × 1.006 = USD 9.054/hr, or 9.054 × 730 = USD 6,609 per month.

Split: (12 × 0.340) + (2 × 0.526) + (2 × 1.006) = 4.080 + 1.052 + 2.012 = USD 7.144/hr, or USD 5,215 per month.

That is a saving of USD 1,394 a month, about 21%. Per request it is small: at 3,200 req/s the split system costs 7.144 / (3,200 × 3,600) = USD 0.00000062 per request, versus USD 0.00000079 for the monolith. Under a tenth of a cent per thousand requests either way.

If your only argument for microservices is the cloud bill, you probably do not have an argument. The real prize is that the LLM replica count and the classifier replica count stopped being the same number.

Look at what else changed. The monolith held 17.9 GB of weights on each of nine cards — 161 GB of GPU memory storing nine copies of a model handling one request every two seconds. The split system holds two copies. And the text service, which carries 95% of the traffic, no longer touches a GPU at all, so it can scale on a spot-instance pool at a third of the price and recover from an eviction in eight seconds instead of ninety.

The bill you pay for splitting

Two costs are unavoidable, and both are quantifiable.

Availability multiplies down. If a request must traverse three services in series and each is independently available 99.9% of the time, the end-to-end availability is 0.9993=0.997000.999^3 = 0.99700, or 99.70%. In a 30-day month that is 0.3% × 43,200 minutes = 130 minutes of downtime, against 43 minutes for a single 99.9% component. Every hop you add is another factor.

Latency adds up per hop. Here is the same logical request, budgeted stage by stage:

StageMonolithFour-service chain
Gateway / HTTP parse1.0 ms1.0 ms
Network hops (in-VPC, p50 1.2 ms each)0.2 ms (4 function calls)4.8 ms (4 hops)
Serialise / deserialise payloads0 ms3.6 ms (588 KiB tensor, ×3 boundaries)
Preprocess8.0 ms8.0 ms
Model inference24.0 ms24.0 ms
Postprocess + response4.0 ms4.0 ms
p50 total37.2 ms45.4 ms
p99 total (hops at 9 ms)~41 ms~76 ms

An extra 8 ms at p50 is usually fine. An extra 35 ms at p99 may not be, if your budget is 100 ms. Notice that the serialisation row — 3.6 ms — is entirely self-inflicted and comes from choosing JSON for tensor payloads. Switching those three boundaries to raw binary over gRPC drops it to about 0.6 ms. This is where people get it wrong: they blame "microservice overhead" for a cost that is really a serialisation format choice.

The organisational cost

Coordination cost grows with the square of the team. A team of n people has n(n−1)/2n(n-1)/2 communication paths: 6 paths at 4 people, 45 at 10, 190 at 20. Microservices are, historically, a way of cutting that graph — each service is owned by one small team that can deploy without asking anyone. If you have four engineers and eleven services, you have inverted the trade: you now pay all the distributed-systems cost and get none of the organisational benefit.

The middle ground people actually ship

Modular monolith

One deployable unit, but with hard internal boundaries: each capability lives behind an interface, owns its own tables, and is forbidden from reaching into another module's internals. Enforce it with import rules in CI, not with good intentions.

Text
moderation/  api/            ← HTTP layer only; may import *_service interfaces  text/    service.py    ← public: classify(text) -> Verdict    model.py      ← private  image/    service.py    ← public: classify(image_bytes) -> Verdict    model.py      ← private  appeals/    service.py    ← public: explain(case) -> strCI rule: no module may import another module's model.py

The payoff is that extracting appeals into its own service later is a two-day job, not a two-quarter rewrite, because the seam already exists and nothing crosses it except a typed call.

Monolith with worker services

The most common and most underrated shape. The API stays monolithic and synchronous; anything slow or GPU-hungry is pushed onto a queue and handled by a separate worker fleet that scales independently.

Text
  client ─► API monolith ─► queue ─► GPU worker pool (scales 0→40)   ▲            │                          │   └── 202 ─────┘                          ▼       + job id                       result store   poll / webhook ◄─────────────────────────┘

This buys you the single biggest microservice benefit — the expensive component scales on its own axis — for a fraction of the complexity. You now have exactly two deployables, not eleven.

API-driven monolith

One process, but every capability is exposed as a clean versioned HTTP resource from day one. Nothing is split, but the contract that a split would need already exists and is already tested by clients.

Choosing, with a decision framework

SignalPoints to monolithPoints to microservices
Team sizeFewer than ~8 engineers totalMultiple teams that must deploy independently
Traffic ratio between capabilitiesWithin ~3× of each otherDiffer by 10× or more
Hardware needsAll models fit the same instance typeMix of CPU-only, small GPU, large GPU
Latency budgetTight — under ~50 ms end to endLoose, or work is already asynchronous
Release cadenceEverything ships together anywayOne model retrains weekly, the rest quarterly
Dependency conflictsOne environment satisfies everythingIrreconcilable CUDA / framework versions
Failure isolationCapabilities fail together acceptablyOne capability must survive another's crash
Operational maturityNo tracing, no service mesh, no on-call rotaDistributed tracing and per-service SLOs already in place

Count the rows. If the right-hand column wins on fewer than four, stay monolithic and make it modular. The default answer for a new AI product is modular monolith plus a worker pool, and you should need evidence to depart from it.

Failure modes with names

  • The distributed monolith. Eleven services that must all be deployed together because they share a database schema and call each other synchronously in a chain. You have every cost of both models and the benefits of neither. The tell: a release checklist that names more than one service.
  • Chatty tensor passing. Splitting preprocessing from inference so that a 588 KiB tensor crosses a network boundary — base64-encoded inside JSON — on every request. Either keep them together or send raw bytes over gRPC.
  • Nine copies of the whale. Scaling a process to satisfy its busiest model while its largest model rides along. This is the opening scenario, and it is the single most expensive mistake in AI serving.
  • Retry storms across a split. Once calls cross the network they can time out, so people add retries. Three services each retrying three times turns one client request into up to 27 downstream calls during a slowdown, converting a brownout into an outage. Retries need budgets and circuit breakers, which a monolith never needed.
  • Splitting before the seams are known. Drawing service boundaries in week two, before you know which parts change together. Boundaries drawn wrong are far more expensive than no boundaries, because now every wrong cut is a network call and a deploy dependency.

The same model, both ways

Consider serving a Hugging Face transformer for a summarisation feature. The monolithic version is a FastAPI app that calls pipeline("summarization") at startup and holds it in a module-level variable. It is roughly 40 lines, deploys as one image, and is the right answer for a product with 20 req/s and one team.

Python
# Monolithic: model lives inside the API processfrom fastapi import FastAPIfrom transformers import pipelineapp = FastAPI()summariser = pipeline("summarization", model="facebook/bart-large-cnn", device=0)@app.post("/v1/summaries")def summarise(body: dict):    out = summariser(body["text"], max_length=120, min_length=30)    return {"summary": out[0]["summary_text"]}

The microservice version keeps the same endpoint but forwards to a dedicated inference server that does dynamic batching and can be scaled and re-versioned on its own:

Python
# Microservice: API is a thin client of a separate inference serviceimport httpxfrom fastapi import FastAPIapp = FastAPI()client = httpx.AsyncClient(base_url="http://summariser.infer.svc", timeout=5.0)@app.post("/v1/summaries")async def summarise(body: dict):    r = await client.post("/predict", json={"text": body["text"], "max_length": 120})    r.raise_for_status()    return {"summary": r.json()["summary"]}

The second version is not better. It is better if the summariser needs a different GPU from the rest of the app, or retrains on a different cadence, or is called by three other products. Otherwise you have added a network hop, a timeout, a retry policy, a second deploy pipeline and a second on-call surface, in exchange for nothing.

What to do on Monday

Start by writing down, for each capability in your system, four numbers: peak requests per second, memory footprint of its weights, the hardware class it needs, and how often it is retrained. Put them in a table. The split almost always announces itself — you will see one row whose traffic is two orders of magnitude off the others, or one row that needs an A100 while everything else needs four vCPUs. Those rows are your first services. Everything else stays together.

Then build the monolith with enforced module boundaries and put the slowest capability behind a queue. Instrument every module boundary with a timer from the first day, because when the time comes to split, the argument will be won or lost on whether you can show that image.classify is 78% of p99 latency and 90% of GPU cost. Architecture decisions made with measurements survive; ones made with preferences get rewritten by the next team.

And set an explicit trigger for the split rather than agonising over it: "we extract the LLM when appeal volume exceeds 20 req/s, or when its retrain cadence goes weekly, whichever comes first". A written trigger turns an argument into a checklist.