Course Content
Building AI Features in Python Backends
5 sections · 23 lessons
Repair and retry, with a limit
Validation turns a bad reply into a known failure. Now you have to decide what to do with it. Throwing away 1.1% of messages is not acceptable, but calling the model again until it gets it right is worse: it has no upper bound on cost or time, and some messages will never pass.
The middle path is a repair: send the model its own reply and the list of validation errors, and ask for a corrected version once. ShipFast's measurements show why once is the right number. On shadow traffic, 1.1% of extraction replies failed the first time. One repair fixed about three quarters of those, leaving 0.28%. A second repair fixed only 0.04 percentage points more, while adding a full call's cost and 1.5 seconds to every message that needed it. One repair, then stop.
Repair or retry?
Repair (show the error)
- Adds the bad reply and the error list to the conversation
- The model knows exactly what to fix
- Costs more input tokens, since the history grows
- Best for value errors: wrong format, date out of range, invented ID
Retry (start again)
- Sends the original request unchanged
- Relies on the model answering differently by chance
- Same cost as the first call
- Best for transport errors: timeouts, 529, network
For validation failures, repair wins: the model gets told "tracking_id: does not appear in the customer's message" and usually fixes exactly that. For transport failures, there is nothing to show the model, so a plain retry with backoff is right, and that belongs in Section 4's resilience layer, not here.
call_structured()
All of ShipFast's structured calls go through one function. It takes a Pydantic model class and returns a validated instance of it, or raises.
1# shipfast/structured.py2from dataclasses import dataclass3from typing import Any, Generic, TypeVar45from pydantic import BaseModel, ValidationError67from shipfast.budget import RequestBudget8from shipfast.llm import LLM, LLMResult910T = TypeVar("T", bound=BaseModel)111213class StructuredOutputError(Exception):14 def __init__(self, reason: str, calls: list[LLMResult]) -> None:15 super().__init__(reason)16 self.calls = calls171819@dataclass20class Structured(Generic[T]):21 value: T22 calls: list[LLMResult]232425def short_errors(err: ValidationError) -> str:26 return "\n".join(f"- {'.'.join(map(str, e['loc'])) or 'reply'}: {e['msg']}"27 for e in err.errors()[:5])Structured carries the value and every LLMResult that produced it, so the caller can add up cost. StructuredOutputError carries the calls too, because a failed extraction still cost money and must still be counted. short_errors turns Pydantic's error list into a few short lines, the same text you saw in the last lesson. LLM is a small Protocol describing anything with a complete() method and a model name; Section 4 explains it. For now, read it as "an LLMClient".
1# shipfast/structured.py (continued)2async def call_structured(llm: LLM, *, feature: str, system: str, user_text: str,3 model_cls: type[T], max_tokens: int = 400, max_attempts: int = 2,4 context: dict[str, Any] | None = None,5 budget: RequestBudget | None = None) -> Structured[T]:6 messages: list[dict] = [{"role": "user", "content": user_text}]7 calls: list[LLMResult] = []8 for attempt in range(1, max_attempts + 1):9 if budget:10 budget.check(llm.model, system + str(messages), max_tokens)11 result = await llm.complete(feature=feature, system=system, messages=messages,12 max_tokens=max_tokens, schema=model_cls.model_json_schema())13 calls.append(result)14 if budget:15 budget.charge(result)16 if result.stop_reason == "max_tokens":17 raise StructuredOutputError("reply cut off at max_tokens", calls)18 try:19 return Structured(model_cls.model_validate_json(result.text, context=context), calls)20 except ValidationError as err:21 problems = short_errors(err)22 messages += [23 {"role": "assistant", "content": result.text},24 {"role": "user", "content": f"Your JSON failed validation:\n{problems}\n"25 "Return the corrected JSON only. Use null if unsure."},26 ]27 raise StructuredOutputError(f"invalid after {max_attempts} attempts: {problems}", calls)max_attempts=2 means one first try and one repair. Read the loop once more for the decisions it makes.
The budget is checked before every attempt, including the repair, whose input is larger because it contains the first reply. A repair that would break the request's budget never starts.
A truncated reply is not repaired. If the model hit max_tokens, sending it back with the same limit usually truncates again, at higher cost. It fails immediately, and the fix is a code change: a higher limit or a smaller schema.
The repair message offers a way out. "Use null if unsure" matters. Without it, a model told "tracking_id does not appear in the message" may invent a different ID. With it, the honest answer, null, is allowed.
The conversation ends with a user turn. The bad reply goes in as an assistant turn and the correction request as a user turn, which is the order every provider expects.
What a failure costs
Say an extraction call uses 470 input tokens and 85 output tokens: $0.0045 on ShipFast's main model. The repair sends the original message, the 85-token bad reply and about 60 tokens of error text, so roughly 620 input tokens and another 85 output: $0.0052. A message that needs a repair costs about 2.2 times a normal one. At 1.1% of messages, that adds about 1.3% to extraction cost. A second repair would add another $0.006 for each of those messages, to fix almost none of them.
After the last attempt
When call_structured raises, the caller decides. In ShipFast, a failed classification becomes unknown and goes to the human triage queue; a failed extraction leaves the fields empty, and the routing rules send the message to a person. The failure is logged with the error list, so the team can see which fields fail most and fix the prompt. Nobody writes unvalidated data anywhere.
Check your understanding
0 of 3 answered
1.A reply fails validation because window_start is "6pm". What should call_structured do on the first attempt?
2.Why does call_structured not repair a reply that stopped at max_tokens?
3.Why does StructuredOutputError carry the list of calls?