Course Content
Building AI Features in Python Backends
5 sections · 23 lessons
A thin client you control
After the last lesson you could call a model from anywhere in ShipFast. That is the problem. If every module calls the SDK directly, then every module picks its own timeout, catches its own set of exceptions, logs differently or not at all, and hard-codes a model name. When you want to change the model, add a fallback or count cost, you edit twenty files.
The fix is the one you already use for databases and payment gateways: put the dependency behind a small class you own. ShipFast's is called LLMClient. It is about 70 lines. Every model call in the course goes through it, and every later lesson (structured output, retries, caching, budgets) builds on it without changing the code that uses it.
The client is deliberately thin. It is not a framework and it does not hide the model. It does five things: holds the model name, enforces a timeout, maps errors, measures usage and logs one line per call.
Each of those five exists because ShipFast's prototype, which called the SDK from three different modules, got it wrong at least once. As you read the code, look for the failure each line prevents. That is the real test of a wrapper: if you cannot name the incident a line stops, the line probably does not belong.
Errors your service can act on
The SDK raises many exception classes. Your service only needs to make three decisions: wait and try again, give up quietly, or page someone because the code is wrong. So LLMClient translates every SDK error into one of three.
1# shipfast/llm.py (part 1)2import logging3import time4from dataclasses import dataclass5from typing import Any, AsyncIterator, Protocol67import anthropic89from shipfast.config import PRICES_PER_MTOK1011log = logging.getLogger("shipfast.llm")121314class LLMError(Exception):15 """The model call did not produce a usable reply."""161718class LLMTimeout(LLMError):19 """No reply within our deadline."""202122class LLMUnavailable(LLMError):23 """429, 5xx, 529 or a network error: worth trying again later."""242526class LLMBadRequest(LLMError):27 """400, 401, 404: a bug on our side. Retrying will not help."""282930@dataclass(frozen=True)31class LLMResult:32 text: str33 model: str34 input_tokens: int35 output_tokens: int36 latency_ms: int37 stop_reason: str | None3839 @property40 def cost_usd(self) -> float:41 price_in, price_out = PRICES_PER_MTOK[self.model]42 return (self.input_tokens * price_in + self.output_tokens * price_out) / 1_000_000LLMResult is what every caller gets back: the text plus everything you need to reason about the call afterwards. cost_usd uses a price table from shipfast/config.py, which the next lesson fills in. The Protocol import is used in Section 4.
Why three error classes and not one? Because the prototype had one: every call sat inside except Exception, followed by a retry. One Friday, a deploy set the model name to claude-opus5, missing a hyphen. Every call returned 404, and every message was retried five times before being marked failed. The on-call engineer saw "triage failed" in the logs with no cause for 40 minutes. With the mapping, the first 404 becomes LLMBadRequest, nothing retries it, and the alert says what is wrong.
Why return an LLMResult and not a string? The prototype's helper returned str. That is how the classifier sometimes passed half a sentence to the parser: the reply had stopped at max_tokens, and the one field that said so was thrown away inside the helper. Carrying stop_reason, tokens and latency with the text means every caller can check the stop reason, and the budget in the next lesson can charge the exact cost. The class is frozen, so no later layer can quietly change the text after it has been logged and charged.
1# shipfast/llm.py (part 2)2def _translate(feature: str, err: anthropic.APIError) -> LLMError:3 if isinstance(err, anthropic.APITimeoutError):4 return LLMTimeout(f"{feature}: timed out")5 if isinstance(err, anthropic.APIConnectionError):6 return LLMUnavailable(f"{feature}: network error")7 if isinstance(err, anthropic.APIStatusError) and (err.status_code == 429 or err.status_code >= 500):8 return LLMUnavailable(f"{feature}: HTTP {err.status_code}")9 return LLMBadRequest(f"{feature}: {err}")101112def _log(feature: str, r: LLMResult) -> None:13 log.info("llm_call feature=%s model=%s in=%d out=%d ms=%d cost_usd=%.5f stop=%s",14 feature, r.model, r.input_tokens, r.output_tokens, r.latency_ms, r.cost_usd, r.stop_reason)The order in _translate matters. In the SDK, a timeout is a kind of connection error, so it must be checked first. Status 529 means "overloaded" and is included in >= 500. Everything else, such as 400 for a malformed request or 404 for a wrong model name, is our bug and becomes LLMBadRequest.
Get the order wrong and nothing crashes; the labels are just wrong. Every timeout would be reported as "network error", and during the first slow evening an engineer would spend an hour checking DNS and firewall rules for a provider that was simply slow. Notice also what the error messages contain: the feature name and the status, never the prompt or the customer's text. An exception message travels to error trackers and log vendors, so it must not carry anything you would not log.
The client
1# shipfast/llm.py (part 3)2class LLMClient:3 def __init__(self, model: str, timeout_s: float = 10.0, effort: str | None = "low",4 sdk: anthropic.AsyncAnthropic | None = None) -> None:5 self.model, self.effort = model, effort6 # max_retries=0: retries live in one visible place (Section 4), not hidden in the SDK.7 self._sdk = sdk or anthropic.AsyncAnthropic(timeout=timeout_s, max_retries=0)89 async def complete(self, *, feature: str, system: str, messages: list[dict],10 max_tokens: int = 512) -> LLMResult:11 output_config = {"effort": self.effort} if self.effort else anthropic.omit12 start = time.perf_counter()13 try:14 resp = await self._sdk.messages.create(15 model=self.model, max_tokens=max_tokens, system=system,16 messages=messages, output_config=output_config)17 except anthropic.APIError as err:18 raise _translate(feature, err) from err19 result = LLMResult(20 text="".join(b.text for b in resp.content if b.type == "text"),21 model=self.model,22 input_tokens=resp.usage.input_tokens,23 output_tokens=resp.usage.output_tokens,24 latency_ms=round((time.perf_counter() - start) * 1000),25 stop_reason=resp.stop_reason,26 )27 _log(feature, result)28 return resultSix decisions in this class are worth defending, and each one answers a failure ShipFast has already had.
The timeout is short and explicit. The SDK's default timeout is 10 minutes, which is right for long generations and wrong for an API endpoint. ShipFast uses 10 seconds by default and less for classification. A request that has not answered in 10 seconds is better treated as failed. The prototype used 30 seconds, and on the first slow evening a handful of stuck calls each held a worker for the full 30 seconds while new messages queued behind them. A short timeout turns a slow provider into fast, countable failures that the retry layer can handle.
SDK retries are off. By default the SDK quietly retries twice. With a 10-second timeout, a single call could then take 30 seconds and you would not see why in your logs. ShipFast sets max_retries=0 here and adds visible retries with a deadline in Section 4. It also stops retries from multiplying: an early pull request added a retry loop around complete() without knowing the SDK already retried, and one failing call became nine requests.
feature is required. Every call is tagged with what it is for: classify, extract, draft. That one string is what lets you answer "which feature is spending the money?" from the logs. In the prototype's second week, daily spend rose 40% and nobody could say why without reading every prompt. With the tag, the same question later took one query: the draft prompt had grown by 300 tokens in an edit. Making it a required argument, with no default, means nobody can forget it.
Arguments are keyword-only. The * in complete(self, *, feature, system, messages, ...) means every argument must be named. The prototype's helper took (system, text) in that order, and one caller passed them the wrong way round, so the customer's message became the system prompt. With keyword-only arguments, that mistake is a TypeError on the first test run, not a quiet prompt-injection risk in production.
"No setting" is anthropic.omit, not None. When effort is not set, the client leaves output_config out of the request entirely. Passing None instead would send "output_config": null in the JSON body, which is not the same thing as "not set" and is not something to depend on. omit is the SDK's way of saying "do not send this field".
The SDK client is injectable. Passing sdk= lets a test hand in a stub, and creating one client per process (not per request) reuses HTTP connections, which saves 50 to 150 ms of connection setup per call. The prototype built a new SDK client inside each request handler, so every classification paid for a fresh TLS handshake, about a tenth of its total latency, for nothing.
A typical log line looks like this:
llm_call feature=classify model=claude-opus-5 in=413 out=38 ms=912 cost_usd=0.00302 stop=end_turnWhat does not belong in the client
It is tempting to keep adding features to LLMClient until it becomes a framework. Resist that. The client knows how to make one call well. It does not know about prompts, customers, caching, retries or budgets.
Each of those arrives later as a separate layer that uses complete(): call_structured() for JSON and repair in Section 3, the cache and ResilientLLM for retries and fallbacks in Section 4, per-customer quotas in Section 5. Keeping them separate means you can test each one alone, turn one off during an incident, and read the client in two minutes when you are debugging at night. The nine-request retry above is what happens otherwise: when two layers both think retrying is their job, the numbers multiply, and nobody sees it until the bill arrives.
There is one exception you will see in Section 3: complete() gains an optional schema argument, because asking the provider for a specific JSON shape is part of making the call, and each provider spells it differently. That difference belongs inside the adapter, where the rest of the code never has to see it.
A fake for tests
Because every caller depends on complete() and nothing else, a test can replace the client with a small scripted fake. No network, no cost, and the same answer every run.
1# tests/fakes.py2from shipfast.llm import LLMResult345class FakeLLM:6 """Scripted replies per feature. Records every call for assertions."""78 model = "claude-haiku-4-5"910 def __init__(self, replies: dict[str, list[str]]) -> None:11 self.replies = {k: list(v) for k, v in replies.items()}12 self.calls: list[dict] = []1314 async def complete(self, *, feature: str, messages: list[dict], **_: object) -> LLMResult:15 self.calls.append({"feature": feature, "messages": messages})16 text = self.replies[feature].pop(0)17 return LLMResult(text=text, model=self.model, input_tokens=600, output_tokens=40,18 latency_ms=5, stop_reason="end_turn")The fake accepts and ignores extra keyword arguments, so it keeps working when later lessons add parameters to complete(). Scripting replies per feature lets one test drive a whole pipeline: the first classify call gets one reply, the first extract call gets another.
The prototype's tests called the real model. Each CI run cost about $3, took four minutes and failed a few times a week, either because the provider was slow or because a reply was worded differently from the assertion. Developers learned to press "re-run" without reading the failure, which is worse than having no tests. With the fake, the pipeline tests run in under a second, cost nothing and fail only when ShipFast's own code changes. The real model is still tested, but on purpose, in the evaluation suite of Section 5.
Check your understanding
0 of 3 answered
1.Why does LLMClient set max_retries=0 on the SDK?
2.The provider returns HTTP 404 because the model name in configuration has a typo. Which error should LLMClient raise?
3.What is the main benefit of requiring a feature argument on every call?