FastAPI Essentials

Course Content

FastAPI Essentials

1 sections · 32 lessons

Can you explain how you would set up unit tests for a FastAPI application?


What you need to know

Here is an app whose real model cannot load in CI — its constructor raises "no GPU in CI" — and a test file that still covers it fully:

Python
# tests/test_sentiment.pyimport pytestfrom fastapi.testclient import TestClientfrom app.main import Sentiment, app, get_modelclass FakeModel:    def __init__(self, result=None, error=None):        self.result, self.error = result, error    def predict(self, text):        if self.error:            raise self.error        return self.result@pytest.fixturedef client():    app.dependency_overrides[get_model] = lambda: FakeModel(Sentiment(label="positive", score=0.9))    yield TestClient(app)          # no `with`: lifespan, and the real model load, never run    app.dependency_overrides.clear()def test_predicts(client):    r = client.post("/sentiment", json={"text": "Loved the delivery"})    assert r.status_code == 200    assert r.json() == {"label": "positive", "score": 0.9}@pytest.mark.parametrize("body", [{}, {"text": ""}, {"text": "x" * 2001}])def test_rejects_bad_input(client, body):    assert client.post("/sentiment", json=body).status_code == 422def test_timeout_maps_to_504(client):    app.dependency_overrides[get_model] = lambda: FakeModel(error=TimeoutError())    assert client.post("/sentiment", json={"text": "hi"}).status_code == 504
Text
$ pytest -q          # these five tests plus the async test shown below......                                                   [100%]6 passed in 0.15s

What each part does:

  • The fixture installs the override before each test and clears it after, so tests do not leak into each other.
  • FakeModel can return a chosen result or raise a chosen error. That is how you test your error mapping without breaking a real model.
  • Parametrize checks three kinds of bad input in one test.

with TestClient(app) or not?

with TestClient(app) as client: runs the app's lifespan startup and shutdown. Use it when the test needs what startup creates. Here, doing so failed immediately with RuntimeError: no GPU in CI, because startup tried to load the real model. Either skip with (as above) or make startup configurable, for example by reading a settings flag that loads a small model in tests.

Testing async code

Python
import httpx, pytest@pytest.mark.anyioasync def test_async_client():    transport = httpx.ASGITransport(app=app)    async with httpx.AsyncClient(transport=transport, base_url="http://test") as ac:        r = await ac.post("/sentiment", json={"text": "late again"})    assert r.status_code == 200

Use this when the test itself must await something, such as an async database fixture. The anyio pytest plugin comes with FastAPI's dependencies; pytest-asyncio works too.

A real-life example

A RAG team's first test suite called the real embedding API and a real vector database. It took 11 minutes, cost money on every CI run, and failed randomly when the provider was slow. Engineers started skipping it.

They rebuilt it around three overrides: get_embedder returns fixed vectors, get_vector_store is an in-memory list, and current_user returns a test user. The suite now runs 180 tests in 6 seconds on every commit. Answer quality moved into a nightly evaluation job that runs 300 real questions against the real model and tracks the score over time. Unit tests check your code; evals check the model.

Follow-up questions to expect

  • "How do you test auth?" — Override current_user for most tests; keep a few tests that send real tokens (valid, expired, wrong role) to check the 401 and 403 paths.
  • "How do you test a streaming endpoint?" — with client.stream("POST", "/chat", json=...) as r: and read r.iter_lines(), with a fake LLM that yields known tokens.
  • "What about the database?" — Override get_db with a session bound to a test database (often SQLite or a Postgres container), wrapped in a transaction that rolls back after each test.