Course Content
FastAPI Essentials
1 sections · 32 lessons
Explain the asynchronous features of FastAPI
What you need to know
Most work in an AI API is I/O-bound: waiting for an LLM provider, a vector database or Postgres. The CPU is idle during that wait. Async lets one process hold hundreds of these waits at once, because a waiting request costs a small coroutine object, not a whole thread. CPU-bound work, such as tokenising or a local model forward pass, is different: the CPU is busy, and async cannot help.
The three cases, measured
This app has three handlers that each "wait" one second. The script fires 10 requests at each handler at the same time:
1import asyncio, time2from fastapi import FastAPI34app = FastAPI()56@app.get("/async-good")7async def async_good():8 await asyncio.sleep(1) # like awaiting an LLM call with httpx.AsyncClient9 return {"ok": True}1011@app.get("/async-bad")12async def async_bad():13 time.sleep(1) # like requests.post(...) inside async def14 return {"ok": True}1516@app.get("/sync-def")17def sync_def():18 time.sleep(1) # plain def: FastAPI runs it in the threadpool19 return {"ok": True}/async-good 10 requests took 1.0s/async-bad 10 requests took 10.0s/sync-def 10 requests took 1.0s/async-bad is the dangerous one. time.sleep never gives control back, so the loop can do nothing else and the requests run one after another. /sync-def is fine because FastAPI sends plain def handlers to a threadpool (40 threads by default, from the anyio library).
| Your code inside the handler | Write the handler as |
|---|---|
Only awaitable I/O (httpx.AsyncClient, asyncpg, async OpenAI SDK) | async def |
Blocking libraries (requests, sync DB driver, sync SDK) | def |
| Heavy CPU work (local model inference) | def, or await run_in_threadpool(...), or a separate model server |
Streaming is where async shines
A chat endpoint that streams tokens keeps each connection open for 10 to 30 seconds. With async, an open stream is just a paused coroutine. Recent FastAPI (0.135 and later) supports Server-Sent Events directly: you yield from the handler.
1from collections.abc import AsyncIterable2from fastapi.sse import EventSourceResponse34@app.post("/chat", response_class=EventSourceResponse)5async def chat(req: ChatRequest) -> AsyncIterable[str]:6 async for token in llm_stream(req.prompt):7 yield tokenEach yielded value goes out as a data: line (for example data: " refund"), and FastAPI sets Cache-Control: no-cache, adds keep-alive pings, and tells nginx not to buffer. On older versions you return a StreamingResponse wrapping a generator.
A real-life example
A bank's support chatbot calls an LLM API that takes about 3 seconds. The first version used the synchronous requests library inside async def chat. With 20 customers asking at once, the last one waited about 60 seconds, because every call blocked the loop in turn. CPU usage was near zero, which confused the team.
The fix was one line: switch to httpx.AsyncClient and await the call. All 20 customers then got answers in about 3 seconds from the same single worker. The test that catches this: send N concurrent requests and check that total time stays flat instead of growing with N.
Follow-up questions to expect
- "Is async faster?" — Not per request. It gives more concurrency for I/O-bound work. A single request takes the same time.
- "How do you run a PyTorch model in FastAPI?" — Load it once at startup, then call it from a
defhandler or withrun_in_threadpool. For heavy traffic, move it to a model server (vLLM, Triton) and call that server with async HTTP. - "What happens when a client disconnects mid-stream?" — The response generator is cancelled. You can also check
await request.is_disconnected()to stop a long generation early and save tokens.