Course Content
FastAPI Essentials
1 sections · 32 lessons
How would you handle exception handling and custom error responses in FastAPI?
What you need to know
There are four kinds of error, and each has its own tool:
| Kind of error | Example | Tool |
|---|---|---|
| Expected, local | Run id not found | raise HTTPException(404, ...) in the route |
| Expected, shared across routes | LLM provider overloaded, quota exceeded | Custom exception + @app.exception_handler |
| Bad input | Wrong type, missing field | Automatic 422; optionally override RequestValidationError |
| Unexpected | A bug, a ZeroDivisionError | Catch-all handler: log it, return a safe 500 |
1import logging, uuid2from fastapi import FastAPI, HTTPException, Request3from fastapi.exceptions import RequestValidationError4from fastapi.responses import JSONResponse56log = logging.getLogger("api")7app = FastAPI()89class ModelOverloaded(Exception):10 pass1112@app.exception_handler(ModelOverloaded)13async def on_overloaded(request: Request, exc: ModelOverloaded):14 return JSONResponse({"error": {"code": "model_overloaded", "message": "Try again shortly"}},15 status_code=503, headers={"Retry-After": "5"})1617@app.exception_handler(RequestValidationError)18async def on_invalid(request: Request, exc: RequestValidationError):19 fields = [".".join(map(str, e["loc"][1:])) for e in exc.errors()]20 return JSONResponse({"error": {"code": "invalid_request", "fields": fields}}, status_code=422)2122@app.exception_handler(Exception)23async def on_crash(request: Request, exc: Exception):24 error_id = uuid.uuid4().hex[:8]25 log.exception("unhandled error %s", error_id) # full traceback goes to logs only26 return JSONResponse({"error": {"code": "internal", "id": error_id}}, status_code=500)With routes that raise each kind, the real responses are:
GET /runs/r9 -> 404 {'detail': 'Run r9 not found'}POST /ask {"question": "busy"} -> 503 {'error': {'code': 'model_overloaded', ...}} Retry-After: 5POST /ask {"temperature": "hot"} -> 422 {'error': {'code': 'invalid_request', 'fields': ['question', 'temperature']}}POST /ask {"question": "bug"} -> 500 {'error': {'code': 'internal', 'id': '3aaa77ff'}}HTTPException stops the handler immediately, like any exception, so there is no need for return after it. A registered handler receives the request and the exception, and returns any response you like. The catch-all handler is the safety net: the user sees an id they can quote to support, and the engineer finds the full traceback in the logs by searching for that id.
Mapping upstream errors honestly
In an LLM service, many errors come from the provider. Translate them into codes that tell your client what to do:
- Provider 429 (rate limited) → your 429 or 503 with
Retry-After: "retry later". - Provider timeout → 504: "the upstream was too slow".
- Content-policy refusal → 400 or 422 with a clear code: "change the input; retrying will not help".
- Your own bug → 500: "not your fault".
A real-life example
A travel company's itinerary assistant calls an LLM provider. During a holiday sale, the provider starts returning 429s. The first version let the provider's exception escape, so users got a bare 500, and the mobile app showed "Something went wrong" with no retry.
The team added a ProviderRateLimited exception, raised by the LLM client wrapper, with one handler returning 503, a Retry-After header and the code model_overloaded. The app now retries after the given delay and shows "High demand, retrying...". Dashboards separate 5xx caused by the provider from real bugs, so the on-call engineer is not woken for something they cannot fix.
Follow-up questions to expect
- "What is the difference between
HTTPExceptionand a normal exception?" —HTTPExceptioncarries a status code and detail, and FastAPI converts it into a response. A normal exception with no handler becomes a 500. - "Where do validation errors come from?" — Pydantic raises them; FastAPI wraps them in
RequestValidationErrorand its default handler returns 422 with adetaillist. - "Should you return error details from the model or the stack trace?" — Never the stack trace or raw provider message. They can leak prompts, keys or internal hostnames. Log them; return a code and an id.