Course Content
FastAPI Essentials
1 sections · 32 lessons
Discuss the role of middleware in FastAPI.
What you need to know
The simple way: @app.middleware("http")
1import time, uuid2from fastapi import FastAPI, Request34app = FastAPI()56@app.middleware("http")7async def add_request_context(request: Request, call_next):8 request_id = request.headers.get("x-request-id") or uuid.uuid4().hex9 start = time.perf_counter()10 response = await call_next(request) # runs the rest of the app11 response.headers["x-request-id"] = request_id12 response.headers["x-process-time"] = f"{time.perf_counter() - start:.3f}"13 return responseThis uses Starlette's BaseHTTPMiddleware. In current Starlette it passes streams through correctly: an SSE endpoint yielding a token every 0.5 s still arrived at 0.53 s, 1.04 s, 1.54 s and 2.04 s in a real test. But the same test showed x-process-time: 0.000. call_next returns as soon as the response starts, so this timing measures time-to-headers, not the 2-second stream.
The streaming-safe way: pure ASGI middleware
1import logging2log = logging.getLogger("access")34class RequestContext:5 def __init__(self, app):6 self.app = app78 async def __call__(self, scope, receive, send):9 if scope["type"] != "http":10 return await self.app(scope, receive, send)11 request_id = dict(scope["headers"]).get(b"x-request-id") or uuid.uuid4().hex.encode()12 start = time.perf_counter()1314 async def send_with_id(message):15 if message["type"] == "http.response.start":16 message["headers"] = [*message["headers"], (b"x-request-id", request_id)]17 await send(message)1819 try:20 await self.app(scope, receive, send_with_id)21 finally:22 log.info("path=%s id=%s total=%.2fs", scope["path"], request_id.decode(),23 time.perf_counter() - start)2425app.add_middleware(RequestContext)With the same stream, the real log line was path=/stream id=req-7f3a total=2.01s. The ASGI version wraps the whole response, including every streamed chunk, so the timing is honest. It is also lighter, and it keeps context variables (used by some tracing libraries) working, which BaseHTTPMiddleware can break.
Order
A real test with two middlewares printed:
B (added last) inA (added first) in handlerA (added first) outB (added last) outThe last one added is the outermost. Add the request-id and logging layer last, so it sees every request, even those another middleware rejects.
Middleware or dependency?
| Use middleware when | Use a dependency when |
|---|---|
| It applies to every request, including 404s | It applies to some routes or routers |
| It works on raw headers and bytes | It needs parsed, validated parameters |
| Examples: request id, CORS, GZip, access logs | Examples: auth, rate limit per plan, DB session |
Built-in middleware includes CORSMiddleware, GZipMiddleware, TrustedHostMiddleware and HTTPSRedirectMiddleware.
A real-life example
An LLM platform's dashboards showed a median latency of 4 ms for its streaming chat endpoint, while users complained that answers took 20 seconds. The timing came from a BaseHTTPMiddleware like the first example: it measured only the time until the first byte, not the full stream.
The team replaced it with a pure ASGI middleware that records the full duration and the time to first token separately, keyed by request id. Both numbers now appear on the dashboard — 600 ms to first token, 18 s total — and the request id is returned to the client, so a support ticket quoting it leads straight to the logs for that conversation.
Follow-up questions to expect
- "Does middleware run for 404s?" — Yes. It wraps the whole app, including requests that match no route. Dependencies never run for those.
- "Can middleware read the request body?" — It can, but reading it in
BaseHTTPMiddlewareis tricky; for body inspection a pure ASGI middleware that wrapsreceive, or a dependency, is safer. - "How expensive is middleware?" — It runs on every request, including health checks, so keep it cheap: no database calls, no heavy parsing.