Building with LLMs

Deploying via FastAPI or Streamlit


The prototype is excellent. It answers questions about your documentation in about three seconds, the demo goes well, and you are asked to put it in front of the support team on Monday.

You containerise it, run it with four worker processes, and deploy. On Monday morning eleven people open it at once. The service stops responding. Not slowly — completely. Requests queue, the load balancer starts returning 502, and someone messages you asking whether it is down.

The arithmetic is brutal and completely predictable. Each request spends about 3 seconds blocked on a network call to the model provider, doing no work at all. With four synchronous workers, you can serve exactly four requests concurrently, so your throughput ceiling is 4 / 3 = 1.33 requests per second. Eleven people clicking within a few seconds of each other exceed it immediately, and the twelfth request waits behind all of them.

Nothing in that failure is about LLMs specifically. It is about a workload that is almost entirely waiting, deployed as though it were computing. That distinction drives nearly every decision in this lesson.

An LLM service is I/O-bound, not CPU-bound. It spends 95% of each request waiting on somebody else's network. Deploy it accordingly, or you will buy four times the hardware to fix a problem that hardware does not fix.

The API service and the internal toolFastAPI• Async handlers over a waiting workload• Streams tokens as they arrive• Health check thatcalls a real dependency• Many workers, because the work is I/OStreamlit• Whole script rerunson every widget change• State lives in thesession, not the process• Fine for a team of ten, not for traffic• Ship it behind the same container
The choice is about who is calling: other software gets an API, humans in your own company get the script.

FastAPI: the async API service

FastAPI is the right default because it is async-native, which is exactly what a waiting-heavy workload needs. While one request waits on the model, the same worker serves others.

Python
from contextlib import asynccontextmanagerfrom fastapi import FastAPI, HTTPException, Requestfrom pydantic import BaseModel, Fieldimport logging, time, uuidresources = {}@asynccontextmanagerasync def lifespan(app: FastAPI):    # startup    resources["chain"] = build_chain()    resources["redis"] = await create_redis_pool()    logging.info("service ready")    yield    # shutdown    await resources["redis"].aclose()    logging.info("service stopped")app = FastAPI(title="Document Assistant", version="1.0.0", lifespan=lifespan)class AskRequest(BaseModel):    question: str = Field(min_length=1, max_length=2000)    session_id: str | None = Noneclass AskResponse(BaseModel):    answer: str    latency_ms: int    request_id: str@app.post("/ask", response_model=AskResponse)async def ask(req: AskRequest):    started, request_id = time.perf_counter(), str(uuid.uuid4())    try:        answer = await resources["chain"].ainvoke({"question": req.question})    except TimeoutError:        raise HTTPException(status_code=504, detail="Upstream model timed out")    return AskResponse(        answer=answer,        latency_ms=int((time.perf_counter() - started) * 1000),        request_id=request_id,    )

Four things there are load-bearing.

lifespan, not @app.on_event. The event decorators are deprecated. The lifespan context manager expresses startup and shutdown as one function, which means the code that opens a connection pool sits next to the code that closes it — and shutdown actually runs, so you stop leaking connections on every redeploy.

Build expensive objects once, at startup. Constructing the chain, the model client and the Redis pool inside the handler adds tens to hundreds of milliseconds to every request and opens a new connection each time. Build once, store, reuse.

async def plus ainvoke. This is the whole point. Declaring the handler async and then calling a blocking function inside it is worse than not being async at all: the blocking call freezes the event loop, so every concurrent request on that worker stalls, not just this one. If you must call blocking code, push it to a thread pool with await asyncio.to_thread(fn, ...).

Pydantic models on input and output. Validation happens before your code runs — a 3,000-character question is rejected with a clear 422 rather than becoming a 12,000-token bill.

Streaming responses

Non-streaming, a user watches a spinner for four seconds. Streaming, text appears in about 400 ms.

Python
from fastapi.responses import StreamingResponseimport json@app.post("/ask/stream")async def ask_stream(req: AskRequest):    async def generate():        try:            async for chunk in resources["chain"].astream({"question": req.question}):                yield f"data: {json.dumps({'delta': chunk})}\n\n"            yield "data: [DONE]\n\n"        except Exception as exc:            logging.exception("stream failed")            yield f"data: {json.dumps({'error': str(exc)})}\n\n"    return StreamingResponse(        generate(),        media_type="text/event-stream",        headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"},    )

The X-Accel-Buffering: no header is the detail that catches people. Nginx buffers proxied responses by default, so your carefully streamed tokens accumulate in the proxy and arrive all at once — the endpoint works perfectly in local testing and appears completely broken behind the load balancer. The error is also yielded into the stream, because once you have sent a 200 and some bytes, you can no longer change the status code.

Batch endpoints, bounded

Python
import asyncioMAX_BATCH = 20sem = asyncio.Semaphore(8)class BatchRequest(BaseModel):    questions: list[str] = Field(min_length=1, max_length=MAX_BATCH)@app.post("/ask/batch")async def ask_batch(req: BatchRequest):    async def one(q):        async with sem:            try:                return {"question": q, "answer": await resources["chain"].ainvoke({"question": q})}            except Exception as exc:                return {"question": q, "error": type(exc).__name__}    results = await asyncio.gather(*(one(q) for q in req.questions))    return {"results": results,            "succeeded": sum(1 for r in results if "answer" in r),            "failed": sum(1 for r in results if "error" in r)}

Two bounds, two different disasters averted. max_length=MAX_BATCH stops one client submitting 10,000 questions and spending your monthly budget in a minute. The semaphore stops even a legal batch of 20 opening 20 simultaneous upstream calls and rate-limiting you. And per-item error handling means one failure returns one error, not a 500 for the whole batch.

Error handling that does not leak

Python
from fastapi.responses import JSONResponsefrom fastapi.exceptions import RequestValidationError@app.exception_handler(RequestValidationError)async def validation_handler(request: Request, exc: RequestValidationError):    return JSONResponse(status_code=422,                        content={"error": "invalid_request", "detail": exc.errors()})@app.exception_handler(Exception)async def unhandled_handler(request: Request, exc: Exception):    request_id = str(uuid.uuid4())    logging.exception("unhandled error request_id=%s path=%s", request_id, request.url.path)    return JSONResponse(        status_code=500,        content={"error": "internal_error",                 "message": "Something went wrong. Quote this ID to support.",                 "request_id": request_id})

The catch-all handler exists so that stack traces never reach clients. A default traceback can disclose file paths, library versions, database names and occasionally the contents of variables. Log the detail with an ID, return the ID, and users can report a problem you can actually find.

Health checks that mean something

Python
@app.get("/health")async def health():    return {"status": "ok"}          # liveness: is the process alive?@app.get("/ready")async def ready():    checks = {}    try:        await resources["redis"].ping(); checks["redis"] = "ok"    except Exception as exc:        checks["redis"] = f"failed: {exc}"    checks["chain"] = "ok" if resources.get("chain") else "missing"    healthy = all(v == "ok" for v in checks.values())    return JSONResponse(status_code=200 if healthy else 503,                        content={"ready": healthy, "checks": checks})

Keep these separate. Liveness answers "should the orchestrator restart this container?" — it must not depend on external services, or a Redis blip restarts every one of your containers simultaneously. Readiness answers "should traffic be routed here?" and legitimately checks dependencies. Conflating them is how a small dependency wobble becomes a full restart storm.

Streamlit: the internal tool

Streamlit turns a Python script into a web UI with no front-end code. It is excellent for internal tools, demos and anything used by tens of people; it is the wrong tool for a public product.

Python
import streamlit as stst.set_page_config(page_title="Document Assistant", page_icon="📄")@st.cache_resourcedef get_chain():    return build_chain()chain = get_chain()if "messages" not in st.session_state:    st.session_state.messages = []for m in st.session_state.messages:    with st.chat_message(m["role"]):        st.markdown(m["content"])if question := st.chat_input("Ask about the documentation"):    st.session_state.messages.append({"role": "user", "content": question})    with st.chat_message("user"):        st.markdown(question)    with st.chat_message("assistant"):        placeholder, full = st.empty(), ""        for chunk in chain.stream({"question": question}):            full += chunk            placeholder.markdown(full + "▌")        placeholder.markdown(full)    st.session_state.messages.append({"role": "assistant", "content": full})

Two Streamlit-specific facts explain most confusion about it. The entire script re-runs top to bottom on every interaction — every keystroke in a widget, every button press. That is why @st.cache_resource is essential: without it you rebuild the chain and reopen the model client on every rerun. And st.session_state is the only thing that survives a rerun; an ordinary variable is recreated from scratch each time.

FastAPIStreamlit
What you getA JSON APIA working UI
Time to first versionHoursMinutes
Concurrency modelAsync, hundreds of connections per workerOne session per connection, thread-per-session
Realistic scaleThousands of usersTens
AuthenticationYours to build, any schemeBasic; usually put behind a proxy
Custom UIAny front endStreamlit's components only
Right forProducts, mobile apps, other servicesInternal tools, demos, data apps

The common production pattern is both: FastAPI holds the logic and the model calls, and Streamlit is a thin client that calls the API over HTTP. You get a UI in an afternoon without trapping your business logic inside a UI framework.

Deployment for Streamlit is either the hosted Community Cloud — connect a repository, set secrets in the dashboard, done, and appropriate for public demos with no sensitive data — or a container like anything else:

Bash
streamlit run app.py \  --server.port=$PORT \  --server.address=0.0.0.0 \  --server.headless=true \  --browser.gatherUsageStats=false

--server.address=0.0.0.0 is required in a container. The default binds to localhost, which means nothing outside the container can reach it, and the symptom is a container that starts cleanly, passes no health check, and logs nothing wrong.

Containerising it

Bash
FROM python:3.12-slim AS builderWORKDIR /appCOPY requirements.txt .RUN pip install --no-cache-dir --user -r requirements.txtFROM python:3.12-slimWORKDIR /appRUN useradd --create-home --uid 1000 appuserCOPY --from=builder /root/.local /home/appuser/.localCOPY --chown=appuser:appuser . .USER appuserENV PATH=/home/appuser/.local/bin:$PATH \    PYTHONUNBUFFERED=1 \    PYTHONDONTWRITEBYTECODE=1EXPOSE 8000HEALTHCHECK --interval=30s --timeout=5s --start-period=20s \  CMD python -c "import urllib.request;urllib.request.urlopen('http://localhost:8000/health')"CMD ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8000", "--workers", "4"]

Four decisions worth understanding rather than copying.

Multi-stage build. Compilers and build headers stay in the builder stage; the final image carries only the installed packages. Typical saving is 300–600 MB, which is faster pulls and a smaller attack surface.

Non-root user. Containers run as root by default. If your process is compromised, root inside the container is a much better starting position for an attacker than an unprivileged user.

PYTHONUNBUFFERED=1. Without it, Python buffers stdout, so your logs appear in chunks, minutes late, and vanish entirely if the container is killed. This one line is the difference between having logs and thinking you do.

Secrets are absent. No ENV OPENAI_API_KEY=..., no ARG carrying a credential. Anything set at build time is baked into a layer and readable by anyone who can pull the image, including through docker history. Secrets are injected at run time.

Worker count for a waiting workload

The standard advice for CPU-bound services is roughly 2 × cores + 1 workers. That advice is wrong here, and it is what produced the outage at the top.

Async workers hold many in-flight requests each, because each request is mostly idle. On a 2-core container, four uvicorn workers can comfortably hold hundreds of concurrent LLM requests — the ceiling is upstream rate limits and memory, not worker count. Adding workers to fix an async I/O-bound service mostly adds memory usage.

The check that actually matters: is the handler genuinely async all the way down? One blocking call — a synchronous requests.get, a blocking database driver, a CPU-heavy parse — turns an async worker back into a synchronous one, silently, and the throughput ceiling collapses to the worker count again.

Composing the stack

Text
services:  api:    build: .    ports: ["8000:8000"]    environment:      - OPENAI_API_KEY=${OPENAI_API_KEY}      - REDIS_URL=redis://redis:6379/0    depends_on:      redis: { condition: service_healthy }    restart: unless-stopped  redis:    image: redis:7-alpine    command: redis-server --maxmemory 512mb --maxmemory-policy allkeys-lru    healthcheck:      test: ["CMD", "redis-cli", "ping"]      interval: 10s      retries: 5    restart: unless-stopped  ui:    build: { context: ., dockerfile: Dockerfile.streamlit }    ports: ["8501:8501"]    environment:      - API_URL=http://api:8000    depends_on: [api]

${OPENAI_API_KEY} reads from your shell or a local .env — the value is never in the compose file. condition: service_healthy waits for Redis to actually answer rather than merely having started, which removes a whole category of flaky startup failures. And the Redis eviction policy caps memory instead of letting the cache grow until the container is killed.

Where to run it

TargetScales to zeroStreamingWatch out forGood for
Google Cloud RunYesYesCold starts; set min instances if latency mattersMost services — simplest good default
AWS ECS / FargateNoYesMore setup: task defs, ALB, target groupsSteady traffic, AWS-standardised teams
AWS Lambda + API GatewayYesYes, with response streaming29 s default integration timeout; cold starts on big depsSpiky, short work
Azure Container InstancesNoYesThin autoscaling on its ownSimple single-container deployments on Azure
Streamlit Community Cloudn/aYesPublic by default; limited resourcesDemos and public prototypes only

Cloud Run is the shortest path for most LLM services: it takes a container, autoscales on concurrent requests, and scales to zero between bursts.

Bash
gcloud run deploy doc-assistant \  --source . \  --region europe-west2 \  --allow-unauthenticated \  --memory 1Gi \  --cpu 1 \  --concurrency 80 \  --timeout 300 \  --min-instances 1 \  --set-secrets OPENAI_API_KEY=openai-key:latest

Three flags deserve attention. --concurrency 80 tells Cloud Run how many simultaneous requests one instance may hold; for an async I/O-bound service the default of 80 is reasonable, and leaving it at 1 (as CPU-bound advice suggests) would make you pay for 80 times the instances. --timeout 300 must exceed your slowest model call plus retries, or the platform kills requests mid-generation. --min-instances 1 costs a little continuously and removes cold starts, which for a Python image with heavy dependencies can be 5–15 seconds — long enough that the first user of every quiet period thinks the service is broken.

The Lambda row carries the sharpest trap: API Gateway's 29-second default integration timeout. A model call that occasionally takes 35 seconds fails there no matter what you set your Lambda timeout to, and it fails in a way that looks like a bug in your code. The limit is no longer absolute: REST APIs can request a longer timeout through a quota increase (at the cost of a lower account throttle limit), and REST API response streaming allows responses of up to 15 minutes. But you only get either if you know the default exists.

Timeouts must nest, outermost shortest. When the platform gives up before the client does, the user sees an error while three layers keep working — and paying — on a request nobody is waiting for.

Knowing what is happening in production

Structured logs

Python
import json, logging, time, uuidfrom contextvars import ContextVarrequest_id_var: ContextVar[str] = ContextVar("request_id", default="-")class JsonFormatter(logging.Formatter):    def format(self, record):        payload = {            "ts": self.formatTime(record),            "level": record.levelname,            "logger": record.name,            "message": record.getMessage(),            "request_id": request_id_var.get(),        }        if record.exc_info:            payload["exception"] = self.formatException(record.exc_info)        return json.dumps(payload)@app.middleware("http")async def add_request_id(request: Request, call_next):    rid = request.headers.get("X-Request-ID", str(uuid.uuid4()))    request_id_var.set(rid)    started = time.perf_counter()    response = await call_next(request)    logging.info("request complete path=%s status=%s ms=%d",                 request.url.path, response.status_code,                 int((time.perf_counter() - started) * 1000))    response.headers["X-Request-ID"] = rid    return response

JSON logs are queryable; free-text logs are grep-able at best. And the request ID threaded through a ContextVar is what lets you reconstruct one user's journey from a hundred thousand interleaved lines — without it, concurrent requests produce logs you cannot untangle.

Metrics

Python
from prometheus_client import Counter, Histogram, make_asgi_appREQUESTS = Counter("llm_requests_total", "Requests", ["endpoint", "status"])LATENCY  = Histogram("llm_request_seconds", "Latency", ["endpoint"],                     buckets=[0.5, 1, 2, 3, 5, 8, 13, 21, 34])TOKENS   = Counter("llm_tokens_total", "Tokens", ["model", "direction"])app.mount("/metrics", make_asgi_app())

Note the bucket edges. The defaults top out around 10 seconds, which is useless for a workload where 8-second responses are normal — everything lands in the overflow bucket and your p95 becomes meaningless. Set buckets that straddle your real distribution.

The four numbers to alert on: p95 latency (users feel this, not the mean), error rate by type (a 429 spike and a 500 spike need different responses), tokens per minute (your cost, live), and upstream provider errors (so you know it is them before a user tells you).

Where people get this wrong

Blocking calls inside async handlers. The most common and most damaging. It looks async, it is measured as async, and it performs worse than a plain sync service because one stuck call freezes every other request on the worker.

Building the model client per request. Adds latency and connections to every call. Build at startup.

Secrets in the image. Baked into a layer, readable forever, survives the file being deleted in a later layer.

Running as root. Free to fix, and it is the difference between a contained incident and a bad one.

Missing PYTHONUNBUFFERED. Your logs are late, chunked, and lost on crash — exactly when you need them.

Unbounded batch endpoints. One client, one request, your whole month's budget.

Health checks that call the model. Every orchestrator probe becomes a paid API call. At one probe per 30 seconds across ten instances, that is 28,800 calls a day for nothing.

Timeouts that do not nest. Client 10 s, gateway 29 s, application 60 s, model call 90 s. The user's request dies at 10 s while three layers keep working on it. Order them so the outermost is the shortest.

Using Streamlit as a public product. It holds a session per connection and was not designed for thousands of anonymous users.

The deployment checklist that would have saved Monday

Before any LLM service takes real traffic, answer these. Each one corresponds to an outage that has happened to somebody.

  1. Is every path from handler to model call genuinely non-blocking? Trace it. One synchronous HTTP client in a helper undoes the whole design.
  2. What is the throughput ceiling, in requests per second? Compute it from concurrency and average latency, then load-test to confirm. If you cannot state the number, you do not know when you will fall over.
  3. Do the timeouts nest correctly, outermost shortest? Write them down in order: client, proxy, platform, application, upstream.
  4. Is every user-influenced quantity bounded? Question length, batch size, concurrency, retries, upload size.
  5. Are liveness and readiness separate, and does neither call a paid API?
  6. Do secrets arrive at run time only? Check docker history on the built image if you are not certain.
  7. Can you find one user's request in the logs from an ID they quote? If not, every support conversation becomes archaeology.
  8. What happens when the provider is down for ten minutes? Cached answer, degraded reply, or a clear error — decide now, not at the time.

The prototype that failed on Monday was not badly written. It was deployed on a mental model borrowed from CPU-bound web services, where four workers means four cores busy. Here, four workers meant four requests waiting and everyone else queueing. Getting that one distinction right — and then checking that nothing blocks the loop — is most of what separates a demo from a service.