Course Content
Capstone Project: Multimodal Assistant
1 sections · 6 lessons
Reasoning & Structured Output
Ask your assistant: "The bill is USD 84.50. What is a 15% tip, and is that reasonable for average service?"
A plain chat call gives you something like: "A 15% tip on USD 84.50 comes to about USD 12.67, which is a standard amount for average service in the US." Fluent, helpful-sounding, and wrong. Fifteen percent of 84.50 is 12.675, which rounds to 12.68. The model did not calculate; it produced the most plausible-looking sequence of digits. Off by a penny here, off by a factor of ten somewhere else.
Then your code tries to use the answer:
result = json.loads(llm_response)# json.decoder.JSONDecodeError: Expecting value: line 1 column 1 (char 0)Because you asked for JSON and got "Sure! Here's the breakdown you asked for:" followed by a fenced code block, followed by a friendly closing sentence.
Two distinct problems, one root cause. A language model generates plausible text. It does not execute arithmetic, and it does not honour a schema, unless you build machinery that makes it do both. That machinery is what this stage is: a reasoning loop that lets the model stop and call a real tool, and a validation layer that either returns a correctly-shaped object or fails loudly.
Prompting harder reduces malformed output; it never eliminates it. Production systems pair a prompt with code that checks the result and fails immediately, rather than letting a bad response corrupt something three steps downstream.
ReAct: reasoning with the ability to act
ReAct interleaves three kinds of step in a loop — a thought, an action, and the observation that action produced. Written out, one run of the tip question looks like this:
Question: The bill is 84.50. What is a 15% tip, and is that reasonable?Thought: This has two parts. The first is arithmetic I should not guess at.Action: calculatorAction Input: 84.50 * 0.15Observation: 12.675Thought: I have the exact figure. The second part is a judgement I can answer from general knowledge without a tool.Final Answer: A 15% tip on 84.50 is 12.68 (12.675 exactly). In the US that is a standard amount for average service; 18-20% is typical for good service.Compare that with chain-of-thought prompting, which asks the model to show its working inline. Chain-of-thought makes the reasoning visible and does genuinely improve multi-step accuracy, but every step is still the model talking to itself. It cannot fetch today's weather, cannot look up what you said last Tuesday, and cannot perform arithmetic it does not already know how to fake.
| Chain-of-thought | ReAct | |
|---|---|---|
| Shows its reasoning | Yes | Yes |
| Calls real code mid-reasoning | No | Yes |
| Can use information not in training data | No | Yes, via observations |
| LLM calls per question | 1 | 1 per loop iteration, typically 2–5 |
| Latency | ~2 s | ~2 s × iterations |
| Can loop forever | No | Yes — this is the main hazard |
| Best for | Logic, explanation, tasks with everything in the prompt | Arithmetic, search, memory lookup, anything with a side effect |
The tool registry
A tool is a Python callable plus enough metadata for the model to know when to reach for it. Registering tools in one place — rather than hard-coding an if action == "calculator" ladder — means adding a capability is a decorator, and the prompt that lists available tools is generated rather than maintained by hand. When those two drift apart, the model starts calling tools that no longer exist.
1# src/reasoning_engine.py2from dataclasses import dataclass3from typing import Callable4from src.exceptions import ReasoningError5from src.utils import logger678@dataclass9class Tool:10 name: str11 description: str # the model reads this to decide when to use it12 func: Callable[[str], str]131415class ToolRegistry:16 def __init__(self):17 self._tools: dict[str, Tool] = {}1819 def register(self, name: str, description: str):20 def decorator(func: Callable[[str], str]):21 self._tools[name] = Tool(name, description, func)22 logger.info("Registered tool: %s", name)23 return func24 return decorator2526 def describe(self) -> str:27 """The exact text injected into the prompt. Generated, never typed."""28 return "\n".join(29 f"- {t.name}: {t.description}" for t in self._tools.values()30 )3132 def names(self) -> set[str]:33 return set(self._tools)3435 def run(self, name: str, argument: str) -> str:36 if name not in self._tools:37 # A hallucinated tool name is normal, not exceptional. Return38 # it to the model as an observation so it can correct itself.39 return (f"Error: no tool named '{name}'. "40 f"Available tools: {', '.join(sorted(self._tools))}")41 try:42 return str(self._tools[name].func(argument))43 except Exception as e:44 logger.warning("Tool %s failed on %r: %s", name, argument, e)45 return f"Error: tool '{name}' failed: {e}"464748registry = ToolRegistry()495051@registry.register("calculator",52 "Evaluate an arithmetic expression, e.g. '84.50 * 0.15'")53def _calc(expression: str) -> str:54 from src.text_processor import safe_calculate55 return str(safe_calculate(expression))565758@registry.register("current_time", "Get the current date and time. Input ignored.")59def _now(_: str) -> str:60 from datetime import datetime61 return datetime.now().strftime("%Y-%m-%d %H:%M")Notice that run() never raises on a bad tool name or a failing tool. It returns an error string. That is deliberate and it is the single most important design decision in the whole loop. A model that invents weather_lookup when you only registered calculator is not a crash — it is a wrong guess, and feeding the correction back as an observation lets it recover on the next iteration. Raise instead, and a single hallucinated name kills the entire request.
The parsing contract
The prompt specifies an output format, and your parser depends on it. Treat that format as an API contract between two pieces of your own system: loosen the wording of the prompt and the regex silently stops matching.
1import re23REACT_TEMPLATE = """Answer the question using this exact format.45Available tools:6{tools}78Format:9Thought: your reasoning about what to do next10Action: the tool name, exactly as listed above11Action Input: the input to pass to that tool12Observation: (filled in for you -- do not write this yourself)13... repeat Thought/Action/Action Input as needed ...14Thought: I now know the final answer15Final Answer: your answer1617Rules:18- Emit at most ONE Action per response, then stop and wait.19- If you can answer without a tool, go straight to Final Answer.20- Never invent a tool that is not listed above.2122Question: {question}23{scratchpad}"""2425ACTION_RE = re.compile(26 r"Action\s*:\s*(?P<tool>[^\n]+?)\s*\n\s*Action\s+Input\s*:\s*(?P<arg>[^\n]+)",27 re.IGNORECASE,28)29FINAL_RE = re.compile(r"Final\s+Answer\s*:\s*(?P<answer>.+)", re.IGNORECASE | re.DOTALL)The instruction "emit at most ONE Action per response, then stop" matters more than it looks. Without it the model happily writes its own fake Observation: lines and hallucinates the tool's output — it will confidently report that the calculator returned 12.67 without the calculator ever having run. You get all the latency of ReAct and none of the correctness.
The loop, with guards
1class ReActAgent:2 def __init__(self, llm, tools: ToolRegistry, max_iterations: int = 6):3 self.llm = llm4 self.tools = tools5 self.max_iterations = max_iterations67 def run(self, question: str) -> dict:8 scratchpad = ""9 trace: list[dict] = []10 seen_actions: set[tuple[str, str]] = set()1112 for step in range(self.max_iterations):13 prompt = REACT_TEMPLATE.format(14 tools=self.tools.describe(),15 question=question,16 scratchpad=scratchpad,17 )18 response = self.llm.chat(prompt)1920 final = FINAL_RE.search(response)21 action = ACTION_RE.search(response)2223 # Check Final Answer first: a response containing both means24 # the model reasoned to a conclusion in the same breath.25 if final and (not action or final.start() < action.start()):26 answer = final.group("answer").strip()27 logger.info("ReAct finished in %d step(s)", step + 1)28 return {"answer": answer, "trace": trace, "steps": step + 1}2930 if not action:31 # Unparseable. Tell the model exactly what went wrong.32 scratchpad += (33 f"\n{response}\nObservation: Your response did not match "34 "the required format. Reply with either 'Action:' plus "35 "'Action Input:', or 'Final Answer:'.\n"36 )37 continue3839 tool = action.group("tool").strip()40 arg = action.group("arg").strip()4142 if (tool, arg) in seen_actions:43 # Same call twice means the model is stuck in a cycle.44 scratchpad += (45 f"\n{response}\nObservation: You already ran {tool} with "46 f"that exact input. Use the earlier result and give a "47 "Final Answer.\n"48 )49 continue50 seen_actions.add((tool, arg))5152 observation = self.tools.run(tool, arg)53 trace.append({"step": step + 1, "tool": tool,54 "input": arg, "observation": observation})55 logger.debug("step %d: %s(%r) -> %r", step + 1, tool, arg, observation)56 scratchpad += f"\n{response}\nObservation: {observation}\n"5758 raise ReasoningError(59 f"no final answer after {self.max_iterations} iterations; "60 f"trace: {trace}"61 )Why the iteration cap is not optional
Every iteration is a full LLM call, and the scratchpad grows with each one. Say the base prompt is 800 tokens and each iteration appends about 150 tokens of thought, action, and observation. Iteration i therefore sends 800+150i tokens, so a six-iteration run sends 6×800+150×(0+1+2+3+4+5)=4,800+2,250=7,050 tokens in total.
Cost grows quadratically in iterations, because each extra step is both an extra call and a longer one. Ten iterations is 8,000+150×45=14,750 tokens — more than double the six-iteration total for less than double the steps. Now remove the cap entirely and let a model that keeps re-running the same lookup spin for 500 iterations before anyone notices: at roughly 1.8 seconds per call that is 15 minutes, and the token total is over 19 million. A cap of 6 with duplicate-action detection turns that from an incident into a clean ReasoningError in about 11 seconds.
Any loop whose exit condition is decided by a language model needs a hard iteration cap enforced by your code, because the model's judgement is exactly the thing that has already failed when the loop will not end.
The three failure modes, and what each guard catches
| Failure mode | What you see | Guard |
|---|---|---|
| Infinite loop | Same tool, same input, forever; latency climbs and nothing returns | max_iterations plus seen_actions duplicate detection |
| Unparseable action | ACTION_RE misses; agent falls through and never progresses | Feed the format error back as an observation instead of failing |
| Hallucinated tool | Model calls web_search that was never registered | run() returns the list of real tool names as an observation |
| Self-written observations | Model fabricates the tool's output; answer is confidently wrong | "At most ONE Action, then stop"; strip anything after the first action match |
| Premature final answer | Model answers without calling the tool it needed | Tool descriptions that state when to use them, not just what they are |
Structured output: why asking nicely is not enough
Put "RESPOND ONLY IN JSON" in capitals at the top of your prompt and the model will comply — most of the time. Most of the time is the problem. Suppose your prompt gets valid, complete JSON on 94% of calls. That sounds acceptable until you count how many structured calls a single user request makes. A request that classifies intent, extracts entities, plans a response, and formats it makes four:
Roughly one session in five breaks. Now add a single retry, where a failed parse is sent back with the validation error attached. Assuming the retry is independent, the per-call failure rate drops from 0.06 to 0.062=0.0036, so per-call success is 0.9964 and the four-call chain succeeds 0.99644=0.9857 of the time — 98.6%. One retry converts a one-in-five failure rate into roughly one in seventy.
1import json2import re3from pydantic import BaseModel, Field, ValidationError456class AssistantAnswer(BaseModel):7 reasoning: str = Field(..., min_length=1,8 description="Why this answer follows")9 answer: str = Field(..., min_length=1,10 description="The answer itself, no preamble")11 confidence: float = Field(..., ge=0.0, le=1.0)12 sources_used: list[str] = Field(default_factory=list)131415JSON_BLOCK = re.compile(r"```(?:json)?\s*(\{.*?\})\s*```", re.DOTALL)161718class StructuredOutput:19 def __init__(self, llm, max_attempts: int = 3):20 self.llm = llm21 self.max_attempts = max_attempts2223 @staticmethod24 def _extract(text: str) -> str:25 """Pull JSON out of whatever conversational wrapper arrived."""26 fenced = JSON_BLOCK.search(text)27 if fenced:28 return fenced.group(1)29 start, end = text.find("{"), text.rfind("}")30 if start != -1 and end > start:31 return text[start:end + 1]32 return text3334 def generate(self, question: str, schema=AssistantAnswer):35 instruction = (36 f"{question}\n\n"37 "Respond with a single JSON object and nothing else. Schema:\n"38 f"{json.dumps(schema.model_json_schema(), indent=2)}"39 )40 last_error = None4142 for attempt in range(1, self.max_attempts + 1):43 raw = self.llm.chat(instruction)44 try:45 return schema.model_validate_json(self._extract(raw))46 except (ValidationError, ValueError) as e:47 last_error = e48 logger.warning("Structured output attempt %d/%d failed: %s",49 attempt, self.max_attempts, e)50 # The repair prompt is the whole trick: the model is told51 # precisely which field was wrong, not just "try again".52 instruction = (53 f"{instruction}\n\nYour previous reply was rejected:\n"54 f"{raw[:500]}\n\nValidation error:\n{e}\n\n"55 "Return corrected JSON only."56 )5758 raise ReasoningError(59 f"could not obtain valid output after {self.max_attempts} "60 f"attempts: {last_error}"61 )Three mechanisms are stacked here and each catches a different thing. _extract handles the common cosmetic failure — perfectly good JSON wrapped in prose or a fence. Pydantic catches the structural failures: a missing confidence, a confidence of 1.5, a sources_used that arrived as a string instead of a list. And the repair loop handles the rest by telling the model the exact validation error, which is far more effective than repeating the original request verbatim.
Native structured output first, validation always
Most current providers can also enforce a JSON schema on the server side, which pushes that 94% much closer to 100% before your code sees anything. With LangChain it is one call, and the Pydantic model you already wrote is the schema:
1from langchain_openai import ChatOpenAI2from config.settings import settings34structured = ChatOpenAI(model=settings.text_model,5 api_key=settings.openai_api_key6 ).with_structured_output(AssistantAnswer)7result = structured.invoke("Is 12.68 a fair tip on a bill of 84.50?")8print(type(result).__name__, result.confidence) # AssistantAnswer, a floatUse it where your model supports it, and keep StructuredOutput as the fallback for models that do not. Either way, keep the Pydantic validation: a schema-shaped reply can still carry a value your code must reject.
The distinction people miss: json.loads() succeeding does not mean the data is usable. {"answer": "42"} is valid JSON and still crashes code that expects confidence. Schema validation is what turns "it parsed" into "it is safe to use", and it is why ge=0.0, le=1.0 earns its place — a model asked for a confidence will occasionally return 95, meaning percent, and silently poisoning every threshold you built on top of it.
Tests
The whole point of the registry-and-parser design is that most of it tests without a real API key. Use a fake LLM that returns scripted responses.
1# tests/test_reasoning.py2import pytest3from src.reasoning_engine import ReActAgent, ToolRegistry4from src.exceptions import ReasoningError5from src.text_processor import safe_calculate678class ScriptedLLM:9 """Returns queued responses in order. No network, fully deterministic."""10 def __init__(self, responses): self.responses = list(responses); self.calls = 011 def chat(self, prompt, **kw):12 self.calls += 113 return self.responses[min(self.calls - 1, len(self.responses) - 1)]141516@pytest.fixture17def registry():18 r = ToolRegistry()19 r.register("calculator", "arithmetic, e.g. '84.50 * 0.15'")(20 lambda x: str(safe_calculate(x))21 )22 return r232425def test_tool_is_actually_called(registry):26 llm = ScriptedLLM([27 "Thought: I need arithmetic.\nAction: calculator\nAction Input: 84.50 * 0.15",28 "Thought: done.\nFinal Answer: 12.68",29 ])30 result = ReActAgent(llm, registry).run("15% tip on 84.50?")31 assert result["steps"] == 232 assert result["trace"][0]["tool"] == "calculator"33 assert result["trace"][0]["observation"].startswith("12.675")343536def test_hallucinated_tool_does_not_crash(registry):37 obs = registry.run("web_search", "anything")38 assert "no tool named" in obs39 assert "calculator" in obs # tells the model what does exist404142def test_repeated_action_is_broken_out_of(registry):43 stuck = "Thought: again.\nAction: calculator\nAction Input: 1+1"44 with pytest.raises(ReasoningError):45 ReActAgent(ScriptedLLM([stuck]), registry, max_iterations=3).run("q")464748def test_confidence_out_of_range_is_rejected():49 from src.reasoning_engine import AssistantAnswer50 from pydantic import ValidationError51 with pytest.raises(ValidationError):52 AssistantAnswer(reasoning="r", answer="a", confidence=95)Two habits make these tests worth having. Assert on the trace, not just the answer — an agent that returns the right number without ever calling the tool has failed even though the string matches. And always include a test that the loop terminates, because a bug in loop-exit logic does not fail your suite, it hangs it.
When things go wrong here
| Symptom | Cause | Fix |
|---|---|---|
| Agent runs the full iteration budget every time | Model never emits Final Answer: — usually the format block is buried or contradictory | Put the format at the top, show one complete worked example, keep the rules to five lines |
Answer is right but trace is empty | Model wrote its own Observation: and never triggered the parser | Enforce "one Action then stop"; truncate the response at the first action match |
ACTION_RE never matches | Model writes Action: calculator("84.5*0.15") — argument inside the action line | Add a fallback regex for the call-style form, or restate the two-line format with an example |
| Tool called with the whole question as input | Tool description says what it is, not what input it wants | Put a concrete input example in the description: e.g. '84.50 * 0.15' |
| JSON parses but a field is missing | Only json.loads() used; no schema | Validate with a Pydantic model; treat a missing field as a retryable error |
| Confidence values of 85 and 0.85 both appear | Prompt did not state the range and no bound was enforced | Field(ge=0.0, le=1.0) plus "confidence between 0 and 1" in the schema description |
| Retry loop never succeeds | Repair prompt repeats the request without the validation error | Append the rejected text and the exact error message to the retry prompt |
| Latency jumped from 2 s to 20 s | Every question now goes through the full ReAct loop | Route: only questions matching tool-shaped intent enter the loop; plain chat stays a single call |
Acceptance criteria for this stage
- The tip question returns exactly
12.675in the trace observation, and the trace showscalculatorwas called with84.50 * 0.15. - A question needing no tool ("who wrote Hamlet?") returns in one iteration —
result["steps"] == 1. - An agent given a scripted LLM that always repeats one action raises
ReasoningErrorwithinmax_iterationsand never hangs. - Calling a tool name that was never registered produces an observation listing the real tool names, and the request still completes.
StructuredOutput.generate()returns anAssistantAnswerwhoseconfidenceis in [0, 1] for 20 consecutive varied questions.- Feeding
{"answer": "42"}to the validator raises, and the repair prompt that follows names the missing fields. - Adding a new tool requires changes in exactly one place — the decorator — with no edit to the prompt template.
- All reasoning tests pass with a dummy API key and no network, using scripted responses.
What this buys you when you wire everything together
Everything above exists so that the rest of the assistant can treat the model as a component with a contract rather than a source of prose. The orchestrator receives an AssistantAnswer object with a typed confidence field, so it can decide to speak a high-confidence answer aloud and ask for clarification on a low-confidence one — a decision that is impossible when the answer is an unstructured paragraph.
The registry pays off the same way. A memory lookup is a tool. An image description is a tool. A weather API is a tool. Each is one decorated function and one sentence of description, and the model discovers it automatically because the prompt's tool list is generated from the registry. That is why the "return errors as observations" rule matters so much: as the tool count grows from two to twelve, hallucinated names and malformed inputs stop being rare events, and a loop that treats them as recoverable feedback keeps working while one that raises falls over.
The trace is the last piece. When a user reports a wrong answer, you have a step-by-step record of which tool ran, with what input, and what it returned. The difference between "the model got it wrong" and "the calculator got 84.50 * 0.15 and returned 12.675, and the model then rounded it to 12.67 in the final answer" is the difference between a shrug and a five-minute fix.