FastAPI Essentials

Course Content

FastAPI Essentials

1 sections · 32 lessons

How would you handle file uploads in FastAPI?


A document on its way into the RAG indexProxy rejectsbodies over 20 MBStarlettespools it:RAM, then diskCheck type,store underyour own idReturn 202with a doc_idWorker parses,chunks and embedsThe size check in Python runs after the bytes have arrived.
The request only accepts the file; the slow work of parsing and embedding happens after the user already has an answer.

What you need to know

Browsers and clients send files as multipart/form-data: the request body is split into parts, one per file or form field. FastAPI needs the python-multipart package to read it.

Declare the file asWhat happensUse when
file: UploadFileSpooled temp file: memory up to 1 MB, then disk. You get filename, content_type, size and a file objectAlmost always
file: bytesThe whole file is read into memoryTiny files only, such as a small image

Here is a document-upload endpoint for a RAG knowledge base:

Python
import uuidfrom pathlib import Pathfrom typing import Annotatedfrom fastapi import FastAPI, File, Form, HTTPException, UploadFileapp = FastAPI()UPLOAD_DIR = Path("/data/uploads")MAX_BYTES = 20 * 1024 * 1024                          # 20 MBALLOWED = {"application/pdf", "text/plain", "text/markdown"}@app.post("/documents", status_code=202)def upload_document(    file: Annotated[UploadFile, File(description="PDF or text for the knowledge base")],    collection: Annotated[str, Form()] = "default",):    if file.content_type not in ALLOWED:        raise HTTPException(415, f"Unsupported type {file.content_type}")    doc_id = uuid.uuid4().hex                         # never build a path from file.filename    dest = UPLOAD_DIR / f"{doc_id}.bin"    size = 0    with dest.open("wb") as out:        while chunk := file.file.read(1024 * 1024):   # copy 1 MB at a time            size += len(chunk)            if size > MAX_BYTES:                out.close(); dest.unlink()                raise HTTPException(413, "File larger than 20 MB")            out.write(chunk)    # enqueue_ingestion(doc_id, collection): parse, chunk and embed in a worker    return {"doc_id": doc_id, "filename": file.filename, "bytes": size, "status": "queued"}

Real responses:

Text
refund-policy.txt, text/plain, 2,200 bytes -> 202 {'filename': 'refund-policy.txt', 'bytes': 2200, 'status': 'queued'}cat.png, image/png                         -> 415 {'detail': 'Unsupported type image/png'}big.txt, 21 MB                             -> 413 {'detail': 'File larger than 20 MB'}

Why each line is there:

  • The handler is plain def, so the blocking disk writes run in the threadpool, not on the event loop.
  • collection is a Form() field. A request can carry files and form fields together, but not a JSON body at the same time.
  • 202 Accepted means "received, processing later". Parsing and embedding a 300-page PDF can take minutes, which is far too long for one HTTP request.
  • filename and content_type come from the client and can lie. ../../etc/passwd is a valid filename. Use your own id for storage, and check the file's real type (for example its first bytes) if it matters.

Where the size limit really belongs

By the time your handler runs, Starlette has already received the whole upload into its temporary file. The check in the code stops you storing and processing a huge file, but the worker has already spent the bandwidth and disk. The first line of defence is the proxy: client_max_body_size 20m; in nginx, or the equivalent setting on your ingress or API gateway, which rejects the request before it reaches Python. For very large files, skip the API entirely: give the client a pre-signed URL and let it upload straight to object storage such as S3.

A real-life example

A legal-tech startup lets lawyers upload case files to "chat with your documents". The first version accepted file: bytes and ran PDF parsing and embedding inside the request. A lawyer uploaded a 180 MB scanned bundle; the worker's memory spiked, the request timed out after 60 seconds, and the browser retried the upload twice.

The new version uses the endpoint above: UploadFile, a 20 MB proxy limit, and a 202 response with a doc_id. A Celery worker parses and embeds the file, and the front end polls GET /documents/{doc_id} for the status. Files over 20 MB go through pre-signed S3 uploads. Upload requests now finish in about a second, whatever the file size.

Follow-up questions to expect

  • "How do you accept several files?" — files: list[UploadFile]. Starlette also limits the number of files and fields per request (1,000 each by default).
  • "Why not process the file in BackgroundTasks?" — It works for small jobs, but the work is lost if the process restarts and it shares the web worker's CPU. Embedding a large document belongs in a real queue with retries.
  • "How do you stream a download back?" — FileResponse for a file on disk, or a redirect to a pre-signed object-storage URL.