Course Content
FastAPI Essentials
1 sections · 32 lessons
How would you deploy a FastAPI application to a production environment?
What you need to know
Running the server
1# Kubernetes: one process per container, scale with replicas2fastapi run app/main.py --port 80003uvicorn app.main:app --host 0.0.0.0 --port 8000 --proxy-headers --timeout-graceful-shutdown 6045# A single VM: several workers in one container or service6uvicorn app.main:app --host 0.0.0.0 --port 8000 --workers 47gunicorn app.main:app -k uvicorn_worker.UvicornWorker -w 4 -b 0.0.0.0:8000fastapi run (from fastapi[standard]) wraps Uvicorn with production defaults. Uvicorn's --workers runs its own process manager and restarts crashed workers. If you use Gunicorn, the worker class now lives in the separate uvicorn-worker package; the old uvicorn.workers module is deprecated. --proxy-headers makes the app see the real client IP from the load balancer's headers.
Loading the model
1from contextlib import asynccontextmanager2from fastapi import FastAPI34@asynccontextmanager5async def lifespan(app: FastAPI):6 app.state.model = load_model() # once per worker, before any request7 app.state.model.predict("warm-up") # first call is often slow8 yield9 app.state.model.close() # on shutdown1011app = FastAPI(lifespan=lifespan)lifespan replaced the old @app.on_event("startup") handlers, which are deprecated. Uvicorn does not accept requests until the code before yield finishes, so no request ever sees a half-loaded model.
Worker count is a memory decision
For a local model, each worker holds its own copy: 4 workers × a 3 GB model = 12 GB of RAM or VRAM. For GPU models the usual pattern is one worker per GPU, batching inside it, or a dedicated model server (vLLM, Triton) with FastAPI as a thin async gateway. For an I/O-bound LLM gateway, one or two async workers per CPU core already handle hundreds of concurrent calls.
The production checklist
- Probes — a startup probe with enough time for the model to load (loading 3 GB can take 60 s); a cheap liveness probe (
/healthz); a readiness probe (/readyz) that fails while dependencies are down or the pod is draining. - Graceful shutdown — on SIGTERM, stop taking new requests and finish current ones. Set Kubernetes
terminationGracePeriodSecondsand--timeout-graceful-shutdownabove your longest stream, or deploys cut answers mid-sentence. - Proxy settings — body-size limit, read timeouts longer than your longest generation, and buffering off for SSE (FastAPI's
EventSourceResponsesendsX-Accel-Buffering: nofor nginx). - Observability — JSON logs with a request id, metrics for latency, error rate, queue depth, tokens and cost.
- Config and secrets — from environment variables or a secret manager via
pydantic-settings; pinned dependencies; a non-root user in the image; never--reload.
A real-life example
A company deploys a document-summary service with a 3 GB local model on Kubernetes. Version one ran gunicorn -w 8 in each pod "for performance". Each pod needed 24 GB of memory and was OOM-killed under load. Startup took four minutes, the default liveness probe killed pods before the models finished loading, and every deploy cut off summaries that were still streaming.
Version two: one Uvicorn worker per pod and more pods; the model loads in lifespan with a warm-up call; a startup probe allows 5 minutes; readiness checks that the model is loaded; the grace period is 120 seconds, longer than the longest summary. The nginx ingress timeout was raised to 300 seconds, with buffering off for the streaming route. Deploys became invisible to users.
Follow-up questions to expect
- "Uvicorn workers or Kubernetes replicas?" — On Kubernetes, prefer one process per container and let replicas scale, so the orchestrator sees each process's health and memory. Use workers on a single VM.
- "How do you handle a very slow model?" — Put it behind a queue or a dedicated model server, and give the API an async job endpoint (
202+ status polling) instead of holding a 5-minute request. - "Where does HTTPS end?" — Usually at the load balancer or ingress; the app speaks plain HTTP inside the cluster, and
--proxy-headersrestores the original scheme and client IP.