Course Content
FastAPI Essentials
1 sections · 32 lessons
Can you explain how you would set up unit tests for a FastAPI application?
What you need to know
Here is an app whose real model cannot load in CI — its constructor raises "no GPU in CI" — and a test file that still covers it fully:
1# tests/test_sentiment.py2import pytest3from fastapi.testclient import TestClient4from app.main import Sentiment, app, get_model56class FakeModel:7 def __init__(self, result=None, error=None):8 self.result, self.error = result, error9 def predict(self, text):10 if self.error:11 raise self.error12 return self.result1314@pytest.fixture15def client():16 app.dependency_overrides[get_model] = lambda: FakeModel(Sentiment(label="positive", score=0.9))17 yield TestClient(app) # no `with`: lifespan, and the real model load, never run18 app.dependency_overrides.clear()1920def test_predicts(client):21 r = client.post("/sentiment", json={"text": "Loved the delivery"})22 assert r.status_code == 20023 assert r.json() == {"label": "positive", "score": 0.9}2425@pytest.mark.parametrize("body", [{}, {"text": ""}, {"text": "x" * 2001}])26def test_rejects_bad_input(client, body):27 assert client.post("/sentiment", json=body).status_code == 4222829def test_timeout_maps_to_504(client):30 app.dependency_overrides[get_model] = lambda: FakeModel(error=TimeoutError())31 assert client.post("/sentiment", json={"text": "hi"}).status_code == 504$ pytest -q # these five tests plus the async test shown below...... [100%]6 passed in 0.15sWhat each part does:
- The fixture installs the override before each test and clears it after, so tests do not leak into each other.
FakeModelcan return a chosen result or raise a chosen error. That is how you test your error mapping without breaking a real model.- Parametrize checks three kinds of bad input in one test.
with TestClient(app) or not?
with TestClient(app) as client: runs the app's lifespan startup and shutdown. Use it when the test needs what startup creates. Here, doing so failed immediately with RuntimeError: no GPU in CI, because startup tried to load the real model. Either skip with (as above) or make startup configurable, for example by reading a settings flag that loads a small model in tests.
Testing async code
1import httpx, pytest23@pytest.mark.anyio4async def test_async_client():5 transport = httpx.ASGITransport(app=app)6 async with httpx.AsyncClient(transport=transport, base_url="http://test") as ac:7 r = await ac.post("/sentiment", json={"text": "late again"})8 assert r.status_code == 200Use this when the test itself must await something, such as an async database fixture. The anyio pytest plugin comes with FastAPI's dependencies; pytest-asyncio works too.
A real-life example
A RAG team's first test suite called the real embedding API and a real vector database. It took 11 minutes, cost money on every CI run, and failed randomly when the provider was slow. Engineers started skipping it.
They rebuilt it around three overrides: get_embedder returns fixed vectors, get_vector_store is an in-memory list, and current_user returns a test user. The suite now runs 180 tests in 6 seconds on every commit. Answer quality moved into a nightly evaluation job that runs 300 real questions against the real model and tracks the score over time. Unit tests check your code; evals check the model.
Follow-up questions to expect
- "How do you test auth?" — Override
current_userfor most tests; keep a few tests that send real tokens (valid, expired, wrong role) to check the 401 and 403 paths. - "How do you test a streaming endpoint?" —
with client.stream("POST", "/chat", json=...) as r:and readr.iter_lines(), with a fake LLM that yields known tokens. - "What about the database?" — Override
get_dbwith a session bound to a test database (often SQLite or a Postgres container), wrapped in a transaction that rolls back after each test.