Course Content
Building with LLMs
4 sections · 10 lessons
Deploying via FastAPI or Streamlit
The prototype is excellent. It answers questions about your documentation in about three seconds, the demo goes well, and you are asked to put it in front of the support team on Monday.
You containerise it, run it with four worker processes, and deploy. On Monday morning eleven people open it at once. The service stops responding. Not slowly — completely. Requests queue, the load balancer starts returning 502, and someone messages you asking whether it is down.
The arithmetic is brutal and completely predictable. Each request spends about 3 seconds blocked on a network call to the model provider, doing no work at all. With four synchronous workers, you can serve exactly four requests concurrently, so your throughput ceiling is 4 / 3 = 1.33 requests per second. Eleven people clicking within a few seconds of each other exceed it immediately, and the twelfth request waits behind all of them.
Nothing in that failure is about LLMs specifically. It is about a workload that is almost entirely waiting, deployed as though it were computing. That distinction drives nearly every decision in this lesson.
An LLM service is I/O-bound, not CPU-bound. It spends 95% of each request waiting on somebody else's network. Deploy it accordingly, or you will buy four times the hardware to fix a problem that hardware does not fix.
FastAPI: the async API service
FastAPI is the right default because it is async-native, which is exactly what a waiting-heavy workload needs. While one request waits on the model, the same worker serves others.
1from contextlib import asynccontextmanager2from fastapi import FastAPI, HTTPException, Request3from pydantic import BaseModel, Field4import logging, time, uuid56resources = {}78@asynccontextmanager9async def lifespan(app: FastAPI):10 # startup11 resources["chain"] = build_chain()12 resources["redis"] = await create_redis_pool()13 logging.info("service ready")14 yield15 # shutdown16 await resources["redis"].aclose()17 logging.info("service stopped")1819app = FastAPI(title="Document Assistant", version="1.0.0", lifespan=lifespan)2021class AskRequest(BaseModel):22 question: str = Field(min_length=1, max_length=2000)23 session_id: str | None = None2425class AskResponse(BaseModel):26 answer: str27 latency_ms: int28 request_id: str2930@app.post("/ask", response_model=AskResponse)31async def ask(req: AskRequest):32 started, request_id = time.perf_counter(), str(uuid.uuid4())33 try:34 answer = await resources["chain"].ainvoke({"question": req.question})35 except TimeoutError:36 raise HTTPException(status_code=504, detail="Upstream model timed out")37 return AskResponse(38 answer=answer,39 latency_ms=int((time.perf_counter() - started) * 1000),40 request_id=request_id,41 )Four things there are load-bearing.
lifespan, not @app.on_event. The event decorators are deprecated. The lifespan context manager expresses startup and shutdown as one function, which means the code that opens a connection pool sits next to the code that closes it — and shutdown actually runs, so you stop leaking connections on every redeploy.
Build expensive objects once, at startup. Constructing the chain, the model client and the Redis pool inside the handler adds tens to hundreds of milliseconds to every request and opens a new connection each time. Build once, store, reuse.
async def plus ainvoke. This is the whole point. Declaring the handler async and then calling a blocking function inside it is worse than not being async at all: the blocking call freezes the event loop, so every concurrent request on that worker stalls, not just this one. If you must call blocking code, push it to a thread pool with await asyncio.to_thread(fn, ...).
Pydantic models on input and output. Validation happens before your code runs — a 3,000-character question is rejected with a clear 422 rather than becoming a 12,000-token bill.
Streaming responses
Non-streaming, a user watches a spinner for four seconds. Streaming, text appears in about 400 ms.
1from fastapi.responses import StreamingResponse2import json34@app.post("/ask/stream")5async def ask_stream(req: AskRequest):6 async def generate():7 try:8 async for chunk in resources["chain"].astream({"question": req.question}):9 yield f"data: {json.dumps({'delta': chunk})}\n\n"10 yield "data: [DONE]\n\n"11 except Exception as exc:12 logging.exception("stream failed")13 yield f"data: {json.dumps({'error': str(exc)})}\n\n"1415 return StreamingResponse(16 generate(),17 media_type="text/event-stream",18 headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"},19 )The X-Accel-Buffering: no header is the detail that catches people. Nginx buffers proxied responses by default, so your carefully streamed tokens accumulate in the proxy and arrive all at once — the endpoint works perfectly in local testing and appears completely broken behind the load balancer. The error is also yielded into the stream, because once you have sent a 200 and some bytes, you can no longer change the status code.
Batch endpoints, bounded
1import asyncio23MAX_BATCH = 204sem = asyncio.Semaphore(8)56class BatchRequest(BaseModel):7 questions: list[str] = Field(min_length=1, max_length=MAX_BATCH)89@app.post("/ask/batch")10async def ask_batch(req: BatchRequest):11 async def one(q):12 async with sem:13 try:14 return {"question": q, "answer": await resources["chain"].ainvoke({"question": q})}15 except Exception as exc:16 return {"question": q, "error": type(exc).__name__}1718 results = await asyncio.gather(*(one(q) for q in req.questions))19 return {"results": results,20 "succeeded": sum(1 for r in results if "answer" in r),21 "failed": sum(1 for r in results if "error" in r)}Two bounds, two different disasters averted. max_length=MAX_BATCH stops one client submitting 10,000 questions and spending your monthly budget in a minute. The semaphore stops even a legal batch of 20 opening 20 simultaneous upstream calls and rate-limiting you. And per-item error handling means one failure returns one error, not a 500 for the whole batch.
Error handling that does not leak
1from fastapi.responses import JSONResponse2from fastapi.exceptions import RequestValidationError34@app.exception_handler(RequestValidationError)5async def validation_handler(request: Request, exc: RequestValidationError):6 return JSONResponse(status_code=422,7 content={"error": "invalid_request", "detail": exc.errors()})89@app.exception_handler(Exception)10async def unhandled_handler(request: Request, exc: Exception):11 request_id = str(uuid.uuid4())12 logging.exception("unhandled error request_id=%s path=%s", request_id, request.url.path)13 return JSONResponse(14 status_code=500,15 content={"error": "internal_error",16 "message": "Something went wrong. Quote this ID to support.",17 "request_id": request_id})The catch-all handler exists so that stack traces never reach clients. A default traceback can disclose file paths, library versions, database names and occasionally the contents of variables. Log the detail with an ID, return the ID, and users can report a problem you can actually find.
Health checks that mean something
1@app.get("/health")2async def health():3 return {"status": "ok"} # liveness: is the process alive?45@app.get("/ready")6async def ready():7 checks = {}8 try:9 await resources["redis"].ping(); checks["redis"] = "ok"10 except Exception as exc:11 checks["redis"] = f"failed: {exc}"12 checks["chain"] = "ok" if resources.get("chain") else "missing"13 healthy = all(v == "ok" for v in checks.values())14 return JSONResponse(status_code=200 if healthy else 503,15 content={"ready": healthy, "checks": checks})Keep these separate. Liveness answers "should the orchestrator restart this container?" — it must not depend on external services, or a Redis blip restarts every one of your containers simultaneously. Readiness answers "should traffic be routed here?" and legitimately checks dependencies. Conflating them is how a small dependency wobble becomes a full restart storm.
Streamlit: the internal tool
Streamlit turns a Python script into a web UI with no front-end code. It is excellent for internal tools, demos and anything used by tens of people; it is the wrong tool for a public product.
1import streamlit as st23st.set_page_config(page_title="Document Assistant", page_icon="📄")45@st.cache_resource6def get_chain():7 return build_chain()89chain = get_chain()1011if "messages" not in st.session_state:12 st.session_state.messages = []1314for m in st.session_state.messages:15 with st.chat_message(m["role"]):16 st.markdown(m["content"])1718if question := st.chat_input("Ask about the documentation"):19 st.session_state.messages.append({"role": "user", "content": question})20 with st.chat_message("user"):21 st.markdown(question)22 with st.chat_message("assistant"):23 placeholder, full = st.empty(), ""24 for chunk in chain.stream({"question": question}):25 full += chunk26 placeholder.markdown(full + "▌")27 placeholder.markdown(full)28 st.session_state.messages.append({"role": "assistant", "content": full})Two Streamlit-specific facts explain most confusion about it. The entire script re-runs top to bottom on every interaction — every keystroke in a widget, every button press. That is why @st.cache_resource is essential: without it you rebuild the chain and reopen the model client on every rerun. And st.session_state is the only thing that survives a rerun; an ordinary variable is recreated from scratch each time.
| FastAPI | Streamlit | |
|---|---|---|
| What you get | A JSON API | A working UI |
| Time to first version | Hours | Minutes |
| Concurrency model | Async, hundreds of connections per worker | One session per connection, thread-per-session |
| Realistic scale | Thousands of users | Tens |
| Authentication | Yours to build, any scheme | Basic; usually put behind a proxy |
| Custom UI | Any front end | Streamlit's components only |
| Right for | Products, mobile apps, other services | Internal tools, demos, data apps |
The common production pattern is both: FastAPI holds the logic and the model calls, and Streamlit is a thin client that calls the API over HTTP. You get a UI in an afternoon without trapping your business logic inside a UI framework.
Deployment for Streamlit is either the hosted Community Cloud — connect a repository, set secrets in the dashboard, done, and appropriate for public demos with no sensitive data — or a container like anything else:
1streamlit run app.py \2 --server.port=$PORT \3 --server.address=0.0.0.0 \4 --server.headless=true \5 --browser.gatherUsageStats=false--server.address=0.0.0.0 is required in a container. The default binds to localhost, which means nothing outside the container can reach it, and the symptom is a container that starts cleanly, passes no health check, and logs nothing wrong.
Containerising it
1FROM python:3.12-slim AS builder2WORKDIR /app3COPY requirements.txt .4RUN pip install --no-cache-dir --user -r requirements.txt56FROM python:3.12-slim7WORKDIR /app89RUN useradd --create-home --uid 1000 appuser10COPY --from=builder /root/.local /home/appuser/.local11COPY --chown=appuser:appuser . .1213USER appuser14ENV PATH=/home/appuser/.local/bin:$PATH \15 PYTHONUNBUFFERED=1 \16 PYTHONDONTWRITEBYTECODE=11718EXPOSE 800019HEALTHCHECK --interval=30s --timeout=5s --start-period=20s \20 CMD python -c "import urllib.request;urllib.request.urlopen('http://localhost:8000/health')"2122CMD ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8000", "--workers", "4"]Four decisions worth understanding rather than copying.
Multi-stage build. Compilers and build headers stay in the builder stage; the final image carries only the installed packages. Typical saving is 300–600 MB, which is faster pulls and a smaller attack surface.
Non-root user. Containers run as root by default. If your process is compromised, root inside the container is a much better starting position for an attacker than an unprivileged user.
PYTHONUNBUFFERED=1. Without it, Python buffers stdout, so your logs appear in chunks, minutes late, and vanish entirely if the container is killed. This one line is the difference between having logs and thinking you do.
Secrets are absent. No ENV OPENAI_API_KEY=..., no ARG carrying a credential. Anything set at build time is baked into a layer and readable by anyone who can pull the image, including through docker history. Secrets are injected at run time.
Worker count for a waiting workload
The standard advice for CPU-bound services is roughly 2 × cores + 1 workers. That advice is wrong here, and it is what produced the outage at the top.
Async workers hold many in-flight requests each, because each request is mostly idle. On a 2-core container, four uvicorn workers can comfortably hold hundreds of concurrent LLM requests — the ceiling is upstream rate limits and memory, not worker count. Adding workers to fix an async I/O-bound service mostly adds memory usage.
The check that actually matters: is the handler genuinely async all the way down? One blocking call — a synchronous requests.get, a blocking database driver, a CPU-heavy parse — turns an async worker back into a synchronous one, silently, and the throughput ceiling collapses to the worker count again.
Composing the stack
services: api: build: . ports: ["8000:8000"] environment: - OPENAI_API_KEY=${OPENAI_API_KEY} - REDIS_URL=redis://redis:6379/0 depends_on: redis: { condition: service_healthy } restart: unless-stopped redis: image: redis:7-alpine command: redis-server --maxmemory 512mb --maxmemory-policy allkeys-lru healthcheck: test: ["CMD", "redis-cli", "ping"] interval: 10s retries: 5 restart: unless-stopped ui: build: { context: ., dockerfile: Dockerfile.streamlit } ports: ["8501:8501"] environment: - API_URL=http://api:8000 depends_on: [api]${OPENAI_API_KEY} reads from your shell or a local .env — the value is never in the compose file. condition: service_healthy waits for Redis to actually answer rather than merely having started, which removes a whole category of flaky startup failures. And the Redis eviction policy caps memory instead of letting the cache grow until the container is killed.
Where to run it
| Target | Scales to zero | Streaming | Watch out for | Good for |
|---|---|---|---|---|
| Google Cloud Run | Yes | Yes | Cold starts; set min instances if latency matters | Most services — simplest good default |
| AWS ECS / Fargate | No | Yes | More setup: task defs, ALB, target groups | Steady traffic, AWS-standardised teams |
| AWS Lambda + API Gateway | Yes | Yes, with response streaming | 29 s default integration timeout; cold starts on big deps | Spiky, short work |
| Azure Container Instances | No | Yes | Thin autoscaling on its own | Simple single-container deployments on Azure |
| Streamlit Community Cloud | n/a | Yes | Public by default; limited resources | Demos and public prototypes only |
Cloud Run is the shortest path for most LLM services: it takes a container, autoscales on concurrent requests, and scales to zero between bursts.
1gcloud run deploy doc-assistant \2 --source . \3 --region europe-west2 \4 --allow-unauthenticated \5 --memory 1Gi \6 --cpu 1 \7 --concurrency 80 \8 --timeout 300 \9 --min-instances 1 \10 --set-secrets OPENAI_API_KEY=openai-key:latestThree flags deserve attention. --concurrency 80 tells Cloud Run how many simultaneous requests one instance may hold; for an async I/O-bound service the default of 80 is reasonable, and leaving it at 1 (as CPU-bound advice suggests) would make you pay for 80 times the instances. --timeout 300 must exceed your slowest model call plus retries, or the platform kills requests mid-generation. --min-instances 1 costs a little continuously and removes cold starts, which for a Python image with heavy dependencies can be 5–15 seconds — long enough that the first user of every quiet period thinks the service is broken.
The Lambda row carries the sharpest trap: API Gateway's 29-second default integration timeout. A model call that occasionally takes 35 seconds fails there no matter what you set your Lambda timeout to, and it fails in a way that looks like a bug in your code. The limit is no longer absolute: REST APIs can request a longer timeout through a quota increase (at the cost of a lower account throttle limit), and REST API response streaming allows responses of up to 15 minutes. But you only get either if you know the default exists.
Timeouts must nest, outermost shortest. When the platform gives up before the client does, the user sees an error while three layers keep working — and paying — on a request nobody is waiting for.
Knowing what is happening in production
Structured logs
1import json, logging, time, uuid2from contextvars import ContextVar34request_id_var: ContextVar[str] = ContextVar("request_id", default="-")56class JsonFormatter(logging.Formatter):7 def format(self, record):8 payload = {9 "ts": self.formatTime(record),10 "level": record.levelname,11 "logger": record.name,12 "message": record.getMessage(),13 "request_id": request_id_var.get(),14 }15 if record.exc_info:16 payload["exception"] = self.formatException(record.exc_info)17 return json.dumps(payload)1819@app.middleware("http")20async def add_request_id(request: Request, call_next):21 rid = request.headers.get("X-Request-ID", str(uuid.uuid4()))22 request_id_var.set(rid)23 started = time.perf_counter()24 response = await call_next(request)25 logging.info("request complete path=%s status=%s ms=%d",26 request.url.path, response.status_code,27 int((time.perf_counter() - started) * 1000))28 response.headers["X-Request-ID"] = rid29 return responseJSON logs are queryable; free-text logs are grep-able at best. And the request ID threaded through a ContextVar is what lets you reconstruct one user's journey from a hundred thousand interleaved lines — without it, concurrent requests produce logs you cannot untangle.
Metrics
1from prometheus_client import Counter, Histogram, make_asgi_app23REQUESTS = Counter("llm_requests_total", "Requests", ["endpoint", "status"])4LATENCY = Histogram("llm_request_seconds", "Latency", ["endpoint"],5 buckets=[0.5, 1, 2, 3, 5, 8, 13, 21, 34])6TOKENS = Counter("llm_tokens_total", "Tokens", ["model", "direction"])78app.mount("/metrics", make_asgi_app())Note the bucket edges. The defaults top out around 10 seconds, which is useless for a workload where 8-second responses are normal — everything lands in the overflow bucket and your p95 becomes meaningless. Set buckets that straddle your real distribution.
The four numbers to alert on: p95 latency (users feel this, not the mean), error rate by type (a 429 spike and a 500 spike need different responses), tokens per minute (your cost, live), and upstream provider errors (so you know it is them before a user tells you).
Where people get this wrong
Blocking calls inside async handlers. The most common and most damaging. It looks async, it is measured as async, and it performs worse than a plain sync service because one stuck call freezes every other request on the worker.
Building the model client per request. Adds latency and connections to every call. Build at startup.
Secrets in the image. Baked into a layer, readable forever, survives the file being deleted in a later layer.
Running as root. Free to fix, and it is the difference between a contained incident and a bad one.
Missing PYTHONUNBUFFERED. Your logs are late, chunked, and lost on crash — exactly when you need them.
Unbounded batch endpoints. One client, one request, your whole month's budget.
Health checks that call the model. Every orchestrator probe becomes a paid API call. At one probe per 30 seconds across ten instances, that is 28,800 calls a day for nothing.
Timeouts that do not nest. Client 10 s, gateway 29 s, application 60 s, model call 90 s. The user's request dies at 10 s while three layers keep working on it. Order them so the outermost is the shortest.
Using Streamlit as a public product. It holds a session per connection and was not designed for thousands of anonymous users.
The deployment checklist that would have saved Monday
Before any LLM service takes real traffic, answer these. Each one corresponds to an outage that has happened to somebody.
- Is every path from handler to model call genuinely non-blocking? Trace it. One synchronous HTTP client in a helper undoes the whole design.
- What is the throughput ceiling, in requests per second? Compute it from concurrency and average latency, then load-test to confirm. If you cannot state the number, you do not know when you will fall over.
- Do the timeouts nest correctly, outermost shortest? Write them down in order: client, proxy, platform, application, upstream.
- Is every user-influenced quantity bounded? Question length, batch size, concurrency, retries, upload size.
- Are liveness and readiness separate, and does neither call a paid API?
- Do secrets arrive at run time only? Check
docker historyon the built image if you are not certain. - Can you find one user's request in the logs from an ID they quote? If not, every support conversation becomes archaeology.
- What happens when the provider is down for ten minutes? Cached answer, degraded reply, or a clear error — decide now, not at the time.
The prototype that failed on Monday was not badly written. It was deployed on a mental model borrowed from CPU-bound web services, where four workers means four cores busy. Here, four workers meant four requests waiting and everyone else queueing. Getting that one distinction right — and then checking that nothing blocks the loop — is most of what separates a demo from a service.