FastAPI Essentials

Course Content

FastAPI Essentials

1 sections · 32 lessons

How would you deploy a FastAPI application to a production environment?


A model service on Kubernetes, outside inIngress — TLS, body limit,long timeouts, no SSE bufferingPod — one Uvicorn process, graceful shutdown 120 sLifespan — load and warm the model onceProbes — startup 5 min,liveness cheap, readiness realHandlers — logs and metrics keyed by request id
Worker count multiplies model memory, so on Kubernetes you scale with pods and let each one hold a single copy.

What you need to know

Running the server

Bash
# Kubernetes: one process per container, scale with replicasfastapi run app/main.py --port 8000uvicorn app.main:app --host 0.0.0.0 --port 8000 --proxy-headers --timeout-graceful-shutdown 60# A single VM: several workers in one container or serviceuvicorn app.main:app --host 0.0.0.0 --port 8000 --workers 4gunicorn app.main:app -k uvicorn_worker.UvicornWorker -w 4 -b 0.0.0.0:8000

fastapi run (from fastapi[standard]) wraps Uvicorn with production defaults. Uvicorn's --workers runs its own process manager and restarts crashed workers. If you use Gunicorn, the worker class now lives in the separate uvicorn-worker package; the old uvicorn.workers module is deprecated. --proxy-headers makes the app see the real client IP from the load balancer's headers.

Loading the model

Python
from contextlib import asynccontextmanagerfrom fastapi import FastAPI@asynccontextmanagerasync def lifespan(app: FastAPI):    app.state.model = load_model()          # once per worker, before any request    app.state.model.predict("warm-up")      # first call is often slow    yield    app.state.model.close()                 # on shutdownapp = FastAPI(lifespan=lifespan)

lifespan replaced the old @app.on_event("startup") handlers, which are deprecated. Uvicorn does not accept requests until the code before yield finishes, so no request ever sees a half-loaded model.

Worker count is a memory decision

For a local model, each worker holds its own copy: 4 workers × a 3 GB model = 12 GB of RAM or VRAM. For GPU models the usual pattern is one worker per GPU, batching inside it, or a dedicated model server (vLLM, Triton) with FastAPI as a thin async gateway. For an I/O-bound LLM gateway, one or two async workers per CPU core already handle hundreds of concurrent calls.

The production checklist

  1. Probes — a startup probe with enough time for the model to load (loading 3 GB can take 60 s); a cheap liveness probe (/healthz); a readiness probe (/readyz) that fails while dependencies are down or the pod is draining.
  2. Graceful shutdown — on SIGTERM, stop taking new requests and finish current ones. Set Kubernetes terminationGracePeriodSeconds and --timeout-graceful-shutdown above your longest stream, or deploys cut answers mid-sentence.
  3. Proxy settings — body-size limit, read timeouts longer than your longest generation, and buffering off for SSE (FastAPI's EventSourceResponse sends X-Accel-Buffering: no for nginx).
  4. Observability — JSON logs with a request id, metrics for latency, error rate, queue depth, tokens and cost.
  5. Config and secrets — from environment variables or a secret manager via pydantic-settings; pinned dependencies; a non-root user in the image; never --reload.

A real-life example

A company deploys a document-summary service with a 3 GB local model on Kubernetes. Version one ran gunicorn -w 8 in each pod "for performance". Each pod needed 24 GB of memory and was OOM-killed under load. Startup took four minutes, the default liveness probe killed pods before the models finished loading, and every deploy cut off summaries that were still streaming.

Version two: one Uvicorn worker per pod and more pods; the model loads in lifespan with a warm-up call; a startup probe allows 5 minutes; readiness checks that the model is loaded; the grace period is 120 seconds, longer than the longest summary. The nginx ingress timeout was raised to 300 seconds, with buffering off for the streaming route. Deploys became invisible to users.

Follow-up questions to expect

  • "Uvicorn workers or Kubernetes replicas?" — On Kubernetes, prefer one process per container and let replicas scale, so the orchestrator sees each process's health and memory. Use workers on a single VM.
  • "How do you handle a very slow model?" — Put it behind a queue or a dedicated model server, and give the API an async job endpoint (202 + status polling) instead of holding a 5-minute request.
  • "Where does HTTPS end?" — Usually at the load balancer or ingress; the app speaks plain HTTP inside the cluster, and --proxy-headers restores the original scheme and client IP.