FastAPI Essentials

Course Content

FastAPI Essentials

1 sections · 32 lessons

Discuss the role of middleware in FastAPI.


Middleware wraps the app like an onionroute handlerroutingA — added firstB — added lasttopbottomB sees the request first and the response last; add request-id logging last so it sees everything.
Middleware runs for every request, even 404s, so it suits cross-cutting work and nothing that depends on the matched route.

What you need to know

The simple way: @app.middleware("http")

Python
import time, uuidfrom fastapi import FastAPI, Requestapp = FastAPI()@app.middleware("http")async def add_request_context(request: Request, call_next):    request_id = request.headers.get("x-request-id") or uuid.uuid4().hex    start = time.perf_counter()    response = await call_next(request)          # runs the rest of the app    response.headers["x-request-id"] = request_id    response.headers["x-process-time"] = f"{time.perf_counter() - start:.3f}"    return response

This uses Starlette's BaseHTTPMiddleware. In current Starlette it passes streams through correctly: an SSE endpoint yielding a token every 0.5 s still arrived at 0.53 s, 1.04 s, 1.54 s and 2.04 s in a real test. But the same test showed x-process-time: 0.000. call_next returns as soon as the response starts, so this timing measures time-to-headers, not the 2-second stream.

The streaming-safe way: pure ASGI middleware

Python
import logginglog = logging.getLogger("access")class RequestContext:    def __init__(self, app):        self.app = app    async def __call__(self, scope, receive, send):        if scope["type"] != "http":            return await self.app(scope, receive, send)        request_id = dict(scope["headers"]).get(b"x-request-id") or uuid.uuid4().hex.encode()        start = time.perf_counter()        async def send_with_id(message):            if message["type"] == "http.response.start":                message["headers"] = [*message["headers"], (b"x-request-id", request_id)]            await send(message)        try:            await self.app(scope, receive, send_with_id)        finally:            log.info("path=%s id=%s total=%.2fs", scope["path"], request_id.decode(),                     time.perf_counter() - start)app.add_middleware(RequestContext)

With the same stream, the real log line was path=/stream id=req-7f3a total=2.01s. The ASGI version wraps the whole response, including every streamed chunk, so the timing is honest. It is also lighter, and it keeps context variables (used by some tracing libraries) working, which BaseHTTPMiddleware can break.

Order

A real test with two middlewares printed:

Text
B (added last) inA (added first) in  handlerA (added first) outB (added last) out

The last one added is the outermost. Add the request-id and logging layer last, so it sees every request, even those another middleware rejects.

Middleware or dependency?

Use middleware whenUse a dependency when
It applies to every request, including 404sIt applies to some routes or routers
It works on raw headers and bytesIt needs parsed, validated parameters
Examples: request id, CORS, GZip, access logsExamples: auth, rate limit per plan, DB session

Built-in middleware includes CORSMiddleware, GZipMiddleware, TrustedHostMiddleware and HTTPSRedirectMiddleware.

A real-life example

An LLM platform's dashboards showed a median latency of 4 ms for its streaming chat endpoint, while users complained that answers took 20 seconds. The timing came from a BaseHTTPMiddleware like the first example: it measured only the time until the first byte, not the full stream.

The team replaced it with a pure ASGI middleware that records the full duration and the time to first token separately, keyed by request id. Both numbers now appear on the dashboard — 600 ms to first token, 18 s total — and the request id is returned to the client, so a support ticket quoting it leads straight to the logs for that conversation.

Follow-up questions to expect

  • "Does middleware run for 404s?" — Yes. It wraps the whole app, including requests that match no route. Dependencies never run for those.
  • "Can middleware read the request body?" — It can, but reading it in BaseHTTPMiddleware is tricky; for body inspection a pure ASGI middleware that wraps receive, or a dependency, is safer.
  • "How expensive is middleware?" — It runs on every request, including health checks, so keep it cheap: no database calls, no heavy parsing.