Capstone Project: Multimodal Assistant

Reasoning & Structured Output


Ask your assistant: "The bill is USD 84.50. What is a 15% tip, and is that reasonable for average service?"

A plain chat call gives you something like: "A 15% tip on USD 84.50 comes to about USD 12.67, which is a standard amount for average service in the US." Fluent, helpful-sounding, and wrong. Fifteen percent of 84.50 is 12.675, which rounds to 12.68. The model did not calculate; it produced the most plausible-looking sequence of digits. Off by a penny here, off by a factor of ten somewhere else.

Then your code tries to use the answer:

Python
result = json.loads(llm_response)# json.decoder.JSONDecodeError: Expecting value: line 1 column 1 (char 0)

Because you asked for JSON and got "Sure! Here's the breakdown you asked for:" followed by a fenced code block, followed by a friendly closing sentence.

Two distinct problems, one root cause. A language model generates plausible text. It does not execute arithmetic, and it does not honour a schema, unless you build machinery that makes it do both. That machinery is what this stage is: a reasoning loop that lets the model stop and call a real tool, and a validation layer that either returns a correctly-shaped object or fails loudly.

Prompting harder reduces malformed output; it never eliminates it. Production systems pair a prompt with code that checks the result and fails immediately, rather than letting a bad response corrupt something three steps downstream.

The ReAct loop, and the guard that ends itThoughtActionActioninputObservationFinalansweriteration 1 of 6toolresult appendedGuards: a hard iteration cap, a parse-failure retry, and an unknown-tool error fed back as an observation.
Observation feeds back into Thought, so without a cap a model that keeps re-planning never reaches the final answer.

ReAct: reasoning with the ability to act

ReAct interleaves three kinds of step in a loop — a thought, an action, and the observation that action produced. Written out, one run of the tip question looks like this:

Text
Question: The bill is 84.50. What is a 15% tip, and is that reasonable?Thought: This has two parts. The first is arithmetic I should not guess at.Action: calculatorAction Input: 84.50 * 0.15Observation: 12.675Thought: I have the exact figure. The second part is a judgement I can         answer from general knowledge without a tool.Final Answer: A 15% tip on 84.50 is 12.68 (12.675 exactly). In the US that              is a standard amount for average service; 18-20% is typical              for good service.

Compare that with chain-of-thought prompting, which asks the model to show its working inline. Chain-of-thought makes the reasoning visible and does genuinely improve multi-step accuracy, but every step is still the model talking to itself. It cannot fetch today's weather, cannot look up what you said last Tuesday, and cannot perform arithmetic it does not already know how to fake.

Chain-of-thoughtReAct
Shows its reasoningYesYes
Calls real code mid-reasoningNoYes
Can use information not in training dataNoYes, via observations
LLM calls per question11 per loop iteration, typically 2–5
Latency~2 s~2 s × iterations
Can loop foreverNoYes — this is the main hazard
Best forLogic, explanation, tasks with everything in the promptArithmetic, search, memory lookup, anything with a side effect

The tool registry

A tool is a Python callable plus enough metadata for the model to know when to reach for it. Registering tools in one place — rather than hard-coding an if action == "calculator" ladder — means adding a capability is a decorator, and the prompt that lists available tools is generated rather than maintained by hand. When those two drift apart, the model starts calling tools that no longer exist.

Python
# src/reasoning_engine.pyfrom dataclasses import dataclassfrom typing import Callablefrom src.exceptions import ReasoningErrorfrom src.utils import logger@dataclassclass Tool:    name: str    description: str      # the model reads this to decide when to use it    func: Callable[[str], str]class ToolRegistry:    def __init__(self):        self._tools: dict[str, Tool] = {}    def register(self, name: str, description: str):        def decorator(func: Callable[[str], str]):            self._tools[name] = Tool(name, description, func)            logger.info("Registered tool: %s", name)            return func        return decorator    def describe(self) -> str:        """The exact text injected into the prompt. Generated, never typed."""        return "\n".join(            f"- {t.name}: {t.description}" for t in self._tools.values()        )    def names(self) -> set[str]:        return set(self._tools)    def run(self, name: str, argument: str) -> str:        if name not in self._tools:            # A hallucinated tool name is normal, not exceptional. Return            # it to the model as an observation so it can correct itself.            return (f"Error: no tool named '{name}'. "                    f"Available tools: {', '.join(sorted(self._tools))}")        try:            return str(self._tools[name].func(argument))        except Exception as e:            logger.warning("Tool %s failed on %r: %s", name, argument, e)            return f"Error: tool '{name}' failed: {e}"registry = ToolRegistry()@registry.register("calculator",                   "Evaluate an arithmetic expression, e.g. '84.50 * 0.15'")def _calc(expression: str) -> str:    from src.text_processor import safe_calculate    return str(safe_calculate(expression))@registry.register("current_time", "Get the current date and time. Input ignored.")def _now(_: str) -> str:    from datetime import datetime    return datetime.now().strftime("%Y-%m-%d %H:%M")

Notice that run() never raises on a bad tool name or a failing tool. It returns an error string. That is deliberate and it is the single most important design decision in the whole loop. A model that invents weather_lookup when you only registered calculator is not a crash — it is a wrong guess, and feeding the correction back as an observation lets it recover on the next iteration. Raise instead, and a single hallucinated name kills the entire request.

The parsing contract

The prompt specifies an output format, and your parser depends on it. Treat that format as an API contract between two pieces of your own system: loosen the wording of the prompt and the regex silently stops matching.

Python
import reREACT_TEMPLATE = """Answer the question using this exact format.Available tools:{tools}Format:Thought: your reasoning about what to do nextAction: the tool name, exactly as listed aboveAction Input: the input to pass to that toolObservation: (filled in for you -- do not write this yourself)... repeat Thought/Action/Action Input as needed ...Thought: I now know the final answerFinal Answer: your answerRules:- Emit at most ONE Action per response, then stop and wait.- If you can answer without a tool, go straight to Final Answer.- Never invent a tool that is not listed above.Question: {question}{scratchpad}"""ACTION_RE = re.compile(    r"Action\s*:\s*(?P<tool>[^\n]+?)\s*\n\s*Action\s+Input\s*:\s*(?P<arg>[^\n]+)",    re.IGNORECASE,)FINAL_RE = re.compile(r"Final\s+Answer\s*:\s*(?P<answer>.+)", re.IGNORECASE | re.DOTALL)

The instruction "emit at most ONE Action per response, then stop" matters more than it looks. Without it the model happily writes its own fake Observation: lines and hallucinates the tool's output — it will confidently report that the calculator returned 12.67 without the calculator ever having run. You get all the latency of ReAct and none of the correctness.

The loop, with guards

Python
class ReActAgent:    def __init__(self, llm, tools: ToolRegistry, max_iterations: int = 6):        self.llm = llm        self.tools = tools        self.max_iterations = max_iterations    def run(self, question: str) -> dict:        scratchpad = ""        trace: list[dict] = []        seen_actions: set[tuple[str, str]] = set()        for step in range(self.max_iterations):            prompt = REACT_TEMPLATE.format(                tools=self.tools.describe(),                question=question,                scratchpad=scratchpad,            )            response = self.llm.chat(prompt)            final = FINAL_RE.search(response)            action = ACTION_RE.search(response)            # Check Final Answer first: a response containing both means            # the model reasoned to a conclusion in the same breath.            if final and (not action or final.start() < action.start()):                answer = final.group("answer").strip()                logger.info("ReAct finished in %d step(s)", step + 1)                return {"answer": answer, "trace": trace, "steps": step + 1}            if not action:                # Unparseable. Tell the model exactly what went wrong.                scratchpad += (                    f"\n{response}\nObservation: Your response did not match "                    "the required format. Reply with either 'Action:' plus "                    "'Action Input:', or 'Final Answer:'.\n"                )                continue            tool = action.group("tool").strip()            arg = action.group("arg").strip()            if (tool, arg) in seen_actions:                # Same call twice means the model is stuck in a cycle.                scratchpad += (                    f"\n{response}\nObservation: You already ran {tool} with "                    f"that exact input. Use the earlier result and give a "                    "Final Answer.\n"                )                continue            seen_actions.add((tool, arg))            observation = self.tools.run(tool, arg)            trace.append({"step": step + 1, "tool": tool,                          "input": arg, "observation": observation})            logger.debug("step %d: %s(%r) -> %r", step + 1, tool, arg, observation)            scratchpad += f"\n{response}\nObservation: {observation}\n"        raise ReasoningError(            f"no final answer after {self.max_iterations} iterations; "            f"trace: {trace}"        )

Why the iteration cap is not optional

Every iteration is a full LLM call, and the scratchpad grows with each one. Say the base prompt is 800 tokens and each iteration appends about 150 tokens of thought, action, and observation. Iteration ii therefore sends 800+150i800 + 150i tokens, so a six-iteration run sends 6×800+150×(0+1+2+3+4+5)=4,800+2,250=7,0506 \times 800 + 150 \times (0+1+2+3+4+5) = 4{,}800 + 2{,}250 = 7{,}050 tokens in total.

Cost grows quadratically in iterations, because each extra step is both an extra call and a longer one. Ten iterations is 8,000+150×45=14,7508{,}000 + 150 \times 45 = 14{,}750 tokens — more than double the six-iteration total for less than double the steps. Now remove the cap entirely and let a model that keeps re-running the same lookup spin for 500 iterations before anyone notices: at roughly 1.8 seconds per call that is 15 minutes, and the token total is over 19 million. A cap of 6 with duplicate-action detection turns that from an incident into a clean ReasoningError in about 11 seconds.

Any loop whose exit condition is decided by a language model needs a hard iteration cap enforced by your code, because the model's judgement is exactly the thing that has already failed when the loop will not end.

The three failure modes, and what each guard catches

Failure modeWhat you seeGuard
Infinite loopSame tool, same input, forever; latency climbs and nothing returnsmax_iterations plus seen_actions duplicate detection
Unparseable actionACTION_RE misses; agent falls through and never progressesFeed the format error back as an observation instead of failing
Hallucinated toolModel calls web_search that was never registeredrun() returns the list of real tool names as an observation
Self-written observationsModel fabricates the tool's output; answer is confidently wrong"At most ONE Action, then stop"; strip anything after the first action match
Premature final answerModel answers without calling the tool it neededTool descriptions that state when to use them, not just what they are

Structured output: why asking nicely is not enough

Put "RESPOND ONLY IN JSON" in capitals at the top of your prompt and the model will comply — most of the time. Most of the time is the problem. Suppose your prompt gets valid, complete JSON on 94% of calls. That sounds acceptable until you count how many structured calls a single user request makes. A request that classifies intent, extracts entities, plans a response, and formats it makes four:

P(all four succeed)=0.944=0.781P(\text{all four succeed}) = 0.94^4 = 0.781

Roughly one session in five breaks. Now add a single retry, where a failed parse is sent back with the validation error attached. Assuming the retry is independent, the per-call failure rate drops from 0.06 to 0.062=0.00360.06^2 = 0.0036, so per-call success is 0.9964 and the four-call chain succeeds 0.99644=0.98570.9964^4 = 0.9857 of the time — 98.6%. One retry converts a one-in-five failure rate into roughly one in seventy.

Python
import jsonimport refrom pydantic import BaseModel, Field, ValidationErrorclass AssistantAnswer(BaseModel):    reasoning: str = Field(..., min_length=1,                           description="Why this answer follows")    answer: str = Field(..., min_length=1,                        description="The answer itself, no preamble")    confidence: float = Field(..., ge=0.0, le=1.0)    sources_used: list[str] = Field(default_factory=list)JSON_BLOCK = re.compile(r"```(?:json)?\s*(\{.*?\})\s*```", re.DOTALL)class StructuredOutput:    def __init__(self, llm, max_attempts: int = 3):        self.llm = llm        self.max_attempts = max_attempts    @staticmethod    def _extract(text: str) -> str:        """Pull JSON out of whatever conversational wrapper arrived."""        fenced = JSON_BLOCK.search(text)        if fenced:            return fenced.group(1)        start, end = text.find("{"), text.rfind("}")        if start != -1 and end > start:            return text[start:end + 1]        return text    def generate(self, question: str, schema=AssistantAnswer):        instruction = (            f"{question}\n\n"            "Respond with a single JSON object and nothing else. Schema:\n"            f"{json.dumps(schema.model_json_schema(), indent=2)}"        )        last_error = None        for attempt in range(1, self.max_attempts + 1):            raw = self.llm.chat(instruction)            try:                return schema.model_validate_json(self._extract(raw))            except (ValidationError, ValueError) as e:                last_error = e                logger.warning("Structured output attempt %d/%d failed: %s",                               attempt, self.max_attempts, e)                # The repair prompt is the whole trick: the model is told                # precisely which field was wrong, not just "try again".                instruction = (                    f"{instruction}\n\nYour previous reply was rejected:\n"                    f"{raw[:500]}\n\nValidation error:\n{e}\n\n"                    "Return corrected JSON only."                )        raise ReasoningError(            f"could not obtain valid output after {self.max_attempts} "            f"attempts: {last_error}"        )

Three mechanisms are stacked here and each catches a different thing. _extract handles the common cosmetic failure — perfectly good JSON wrapped in prose or a fence. Pydantic catches the structural failures: a missing confidence, a confidence of 1.5, a sources_used that arrived as a string instead of a list. And the repair loop handles the rest by telling the model the exact validation error, which is far more effective than repeating the original request verbatim.

Native structured output first, validation always

Most current providers can also enforce a JSON schema on the server side, which pushes that 94% much closer to 100% before your code sees anything. With LangChain it is one call, and the Pydantic model you already wrote is the schema:

Python
from langchain_openai import ChatOpenAIfrom config.settings import settingsstructured = ChatOpenAI(model=settings.text_model,                        api_key=settings.openai_api_key                        ).with_structured_output(AssistantAnswer)result = structured.invoke("Is 12.68 a fair tip on a bill of 84.50?")print(type(result).__name__, result.confidence)   # AssistantAnswer, a float

Use it where your model supports it, and keep StructuredOutput as the fallback for models that do not. Either way, keep the Pydantic validation: a schema-shaped reply can still carry a value your code must reject.

The distinction people miss: json.loads() succeeding does not mean the data is usable. {"answer": "42"} is valid JSON and still crashes code that expects confidence. Schema validation is what turns "it parsed" into "it is safe to use", and it is why ge=0.0, le=1.0 earns its place — a model asked for a confidence will occasionally return 95, meaning percent, and silently poisoning every threshold you built on top of it.

Tests

The whole point of the registry-and-parser design is that most of it tests without a real API key. Use a fake LLM that returns scripted responses.

Python
# tests/test_reasoning.pyimport pytestfrom src.reasoning_engine import ReActAgent, ToolRegistryfrom src.exceptions import ReasoningErrorfrom src.text_processor import safe_calculateclass ScriptedLLM:    """Returns queued responses in order. No network, fully deterministic."""    def __init__(self, responses): self.responses = list(responses); self.calls = 0    def chat(self, prompt, **kw):        self.calls += 1        return self.responses[min(self.calls - 1, len(self.responses) - 1)]@pytest.fixturedef registry():    r = ToolRegistry()    r.register("calculator", "arithmetic, e.g. '84.50 * 0.15'")(        lambda x: str(safe_calculate(x))    )    return rdef test_tool_is_actually_called(registry):    llm = ScriptedLLM([        "Thought: I need arithmetic.\nAction: calculator\nAction Input: 84.50 * 0.15",        "Thought: done.\nFinal Answer: 12.68",    ])    result = ReActAgent(llm, registry).run("15% tip on 84.50?")    assert result["steps"] == 2    assert result["trace"][0]["tool"] == "calculator"    assert result["trace"][0]["observation"].startswith("12.675")def test_hallucinated_tool_does_not_crash(registry):    obs = registry.run("web_search", "anything")    assert "no tool named" in obs    assert "calculator" in obs      # tells the model what does existdef test_repeated_action_is_broken_out_of(registry):    stuck = "Thought: again.\nAction: calculator\nAction Input: 1+1"    with pytest.raises(ReasoningError):        ReActAgent(ScriptedLLM([stuck]), registry, max_iterations=3).run("q")def test_confidence_out_of_range_is_rejected():    from src.reasoning_engine import AssistantAnswer    from pydantic import ValidationError    with pytest.raises(ValidationError):        AssistantAnswer(reasoning="r", answer="a", confidence=95)

Two habits make these tests worth having. Assert on the trace, not just the answer — an agent that returns the right number without ever calling the tool has failed even though the string matches. And always include a test that the loop terminates, because a bug in loop-exit logic does not fail your suite, it hangs it.

When things go wrong here

SymptomCauseFix
Agent runs the full iteration budget every timeModel never emits Final Answer: — usually the format block is buried or contradictoryPut the format at the top, show one complete worked example, keep the rules to five lines
Answer is right but trace is emptyModel wrote its own Observation: and never triggered the parserEnforce "one Action then stop"; truncate the response at the first action match
ACTION_RE never matchesModel writes Action: calculator("84.5*0.15") — argument inside the action lineAdd a fallback regex for the call-style form, or restate the two-line format with an example
Tool called with the whole question as inputTool description says what it is, not what input it wantsPut a concrete input example in the description: e.g. '84.50 * 0.15'
JSON parses but a field is missingOnly json.loads() used; no schemaValidate with a Pydantic model; treat a missing field as a retryable error
Confidence values of 85 and 0.85 both appearPrompt did not state the range and no bound was enforcedField(ge=0.0, le=1.0) plus "confidence between 0 and 1" in the schema description
Retry loop never succeedsRepair prompt repeats the request without the validation errorAppend the rejected text and the exact error message to the retry prompt
Latency jumped from 2 s to 20 sEvery question now goes through the full ReAct loopRoute: only questions matching tool-shaped intent enter the loop; plain chat stays a single call

Acceptance criteria for this stage

  1. The tip question returns exactly 12.675 in the trace observation, and the trace shows calculator was called with 84.50 * 0.15.
  2. A question needing no tool ("who wrote Hamlet?") returns in one iteration — result["steps"] == 1.
  3. An agent given a scripted LLM that always repeats one action raises ReasoningError within max_iterations and never hangs.
  4. Calling a tool name that was never registered produces an observation listing the real tool names, and the request still completes.
  5. StructuredOutput.generate() returns an AssistantAnswer whose confidence is in [0, 1] for 20 consecutive varied questions.
  6. Feeding {"answer": "42"} to the validator raises, and the repair prompt that follows names the missing fields.
  7. Adding a new tool requires changes in exactly one place — the decorator — with no edit to the prompt template.
  8. All reasoning tests pass with a dummy API key and no network, using scripted responses.

What this buys you when you wire everything together

Everything above exists so that the rest of the assistant can treat the model as a component with a contract rather than a source of prose. The orchestrator receives an AssistantAnswer object with a typed confidence field, so it can decide to speak a high-confidence answer aloud and ask for clarification on a low-confidence one — a decision that is impossible when the answer is an unstructured paragraph.

The registry pays off the same way. A memory lookup is a tool. An image description is a tool. A weather API is a tool. Each is one decorated function and one sentence of description, and the model discovers it automatically because the prompt's tool list is generated from the registry. That is why the "return errors as observations" rule matters so much: as the tool count grows from two to twelve, hallucinated names and malformed inputs stop being rare events, and a loop that treats them as recoverable feedback keeps working while one that raises falls over.

The trace is the last piece. When a user reports a wrong answer, you have a step-by-step record of which tool ran, with what input, and what it returned. The difference between "the model got it wrong" and "the calculator got 84.50 * 0.15 and returned 12.675, and the model then rounded it to 12.67 in the final answer" is the difference between a shrug and a five-minute fix.