Building AI Features in Python Backends

Repair and retry, with a limit


Validation turns a bad reply into a known failure. Now you have to decide what to do with it. Throwing away 1.1% of messages is not acceptable, but calling the model again until it gets it right is worse: it has no upper bound on cost or time, and some messages will never pass.

The middle path is a repair: send the model its own reply and the list of validation errors, and ask for a corrected version once. ShipFast's measurements show why once is the right number. On shadow traffic, 1.1% of extraction replies failed the first time. One repair fixed about three quarters of those, leaving 0.28%. A second repair fixed only 0.04 percentage points more, while adding a full call's cost and 1.5 seconds to every message that needed it. One repair, then stop.

One repair, then a clean stopcall with schemavalidatesend errorsback oncevalidate againneeds_reviewnullnever a second timecost of bothcalls kept1.1% fail first; one repair leaves 0.28%; a second would fix 0.04 points more.
One repair recovers three quarters of failures, and a second would cost a full call to fix almost none.

Repair or retry?

Repair (show the error)

  • Adds the bad reply and the error list to the conversation
  • The model knows exactly what to fix
  • Costs more input tokens, since the history grows
  • Best for value errors: wrong format, date out of range, invented ID

Retry (start again)

  • Sends the original request unchanged
  • Relies on the model answering differently by chance
  • Same cost as the first call
  • Best for transport errors: timeouts, 529, network

For validation failures, repair wins: the model gets told "tracking_id: does not appear in the customer's message" and usually fixes exactly that. For transport failures, there is nothing to show the model, so a plain retry with backoff is right, and that belongs in Section 4's resilience layer, not here.

call_structured()

All of ShipFast's structured calls go through one function. It takes a Pydantic model class and returns a validated instance of it, or raises.

Python
# shipfast/structured.pyfrom dataclasses import dataclassfrom typing import Any, Generic, TypeVarfrom pydantic import BaseModel, ValidationErrorfrom shipfast.budget import RequestBudgetfrom shipfast.llm import LLM, LLMResultT = TypeVar("T", bound=BaseModel)class StructuredOutputError(Exception):    def __init__(self, reason: str, calls: list[LLMResult]) -> None:        super().__init__(reason)        self.calls = calls@dataclassclass Structured(Generic[T]):    value: T    calls: list[LLMResult]def short_errors(err: ValidationError) -> str:    return "\n".join(f"- {'.'.join(map(str, e['loc'])) or 'reply'}: {e['msg']}"                     for e in err.errors()[:5])

Structured carries the value and every LLMResult that produced it, so the caller can add up cost. StructuredOutputError carries the calls too, because a failed extraction still cost money and must still be counted. short_errors turns Pydantic's error list into a few short lines, the same text you saw in the last lesson. LLM is a small Protocol describing anything with a complete() method and a model name; Section 4 explains it. For now, read it as "an LLMClient".

Python
# shipfast/structured.py (continued)async def call_structured(llm: LLM, *, feature: str, system: str, user_text: str,                          model_cls: type[T], max_tokens: int = 400, max_attempts: int = 2,                          context: dict[str, Any] | None = None,                          budget: RequestBudget | None = None) -> Structured[T]:    messages: list[dict] = [{"role": "user", "content": user_text}]    calls: list[LLMResult] = []    for attempt in range(1, max_attempts + 1):        if budget:            budget.check(llm.model, system + str(messages), max_tokens)        result = await llm.complete(feature=feature, system=system, messages=messages,                                    max_tokens=max_tokens, schema=model_cls.model_json_schema())        calls.append(result)        if budget:            budget.charge(result)        if result.stop_reason == "max_tokens":            raise StructuredOutputError("reply cut off at max_tokens", calls)        try:            return Structured(model_cls.model_validate_json(result.text, context=context), calls)        except ValidationError as err:            problems = short_errors(err)        messages += [            {"role": "assistant", "content": result.text},            {"role": "user", "content": f"Your JSON failed validation:\n{problems}\n"                                        "Return the corrected JSON only. Use null if unsure."},        ]    raise StructuredOutputError(f"invalid after {max_attempts} attempts: {problems}", calls)

max_attempts=2 means one first try and one repair. Read the loop once more for the decisions it makes.

The budget is checked before every attempt, including the repair, whose input is larger because it contains the first reply. A repair that would break the request's budget never starts.

A truncated reply is not repaired. If the model hit max_tokens, sending it back with the same limit usually truncates again, at higher cost. It fails immediately, and the fix is a code change: a higher limit or a smaller schema.

The repair message offers a way out. "Use null if unsure" matters. Without it, a model told "tracking_id does not appear in the message" may invent a different ID. With it, the honest answer, null, is allowed.

The conversation ends with a user turn. The bad reply goes in as an assistant turn and the correction request as a user turn, which is the order every provider expects.

What a failure costs

Say an extraction call uses 470 input tokens and 85 output tokens: $0.0045 on ShipFast's main model. The repair sends the original message, the 85-token bad reply and about 60 tokens of error text, so roughly 620 input tokens and another 85 output: $0.0052. A message that needs a repair costs about 2.2 times a normal one. At 1.1% of messages, that adds about 1.3% to extraction cost. A second repair would add another $0.006 for each of those messages, to fix almost none of them.

After the last attempt

When call_structured raises, the caller decides. In ShipFast, a failed classification becomes unknown and goes to the human triage queue; a failed extraction leaves the fields empty, and the routing rules send the message to a person. The failure is logged with the error list, so the team can see which fields fail most and fix the prompt. Nobody writes unvalidated data anywhere.

Check your understanding

0 of 3 answered

1.A reply fails validation because window_start is "6pm". What should call_structured do on the first attempt?

2.Why does call_structured not repair a reply that stopped at max_tokens?

3.Why does StructuredOutputError carry the list of calls?