- MantraMindAI
- Blog
- AI Agents & Automation
Building an AI agent: the loop, the tools, and the blast radius
Jai Rao
August 22, 202618 min read
An agent is a loop: the model proposes a tool call, your code runs it, the result goes back. The hard parts are tool design, termination, error compounding, and blast radius.
Most writing about AI agents describes them the way a brochure describes an aeroplane. Planning. Reasoning. Autonomy. Goals. Then you open the source of a working agent and find something anticlimactic: a while loop, a dictionary of functions, and about forty lines of glue. The model does not run anything. It emits a structured request — call read_file with this path — and your code decides whether that actually happens.
Closing that gap early is worth the trouble, because once the loop is visible the interesting problems come into focus, and almost none of them are about the model. They are about what your tools hand back, when the loop stops, what happens at step nine when step three was quietly wrong, and what the thing can reach on the day it is confidently doing the wrong job. The loop comes first here, then each of those in turn.
The whole thing is a while loop
The contract is narrow. You send a list of messages and a list of tool definitions. The model replies with either text, meaning it is done, or one or more tool use blocks, meaning it wants something executed. It has no ability to execute anything itself. Your code runs the tool, appends the result to the conversation, and sends the whole thing again. That is the entire mechanism, and here it is in full.
import anthropicclient = anthropic.Anthropic()def run(user_input, tools, handlers, max_steps=12): messages = [{"role": "user", "content": user_input}] for step in range(max_steps): reply = client.messages.create( model="claude-opus-5", max_tokens=4096, tools=tools, messages=messages, ) messages.append({"role": "assistant", "content": reply.content}) calls = [b for b in reply.content if b.type == "tool_use"] if not calls: # it answered instead of acting: done return reply, messages results = [dispatch(handlers, c) for c in calls] messages.append({"role": "user", "content": results}) raise StepBudgetExceeded(max_steps) # your own exception classThree details in there matter more than they look. The model never touches your process, so every capability it has is one you handed it deliberately. The conversation grows monotonically — each step re-sends everything that came before — which turns context into both a cost problem and a correctness problem later on. And the loop is bounded by max_steps, not by the model deciding it has finished. One more thing to get right from the start: a single assistant message can contain several tool use blocks. Execute them, then return all of their results in one user message. Splitting them across separate messages is malformed, and it teaches the model to stop asking for parallel work.
Tools are the only API the model has
The model sees three things about a tool: its name, its description, and a JSON Schema for its arguments. That is the whole interface. It cannot read your code, guess your column names, or find out that limit silently caps at 50. Everything it needs to call the tool correctly has to be in those three fields, which makes tool definitions a writing task as much as an engineering one. A description that says what the tool does is table stakes; the one that says when to call it is what actually changes behaviour.
SEARCH_ORDERS = { "name": "search_orders", "description": ( "Find orders belonging to one customer. Call this when the user asks " "about order status, refunds, or delivery for a specific customer. " "Returns at most 20 orders, newest first, with a row count." ), "input_schema": { "type": "object", "properties": { "customer_id": {"type": "string", "description": "Internal id, e.g. cus_8471"}, "status": {"type": "string", "enum": ["open", "shipped", "cancelled"]}, "limit": {"type": "integer", "minimum": 1, "maximum": 20, "default": 10}, }, "required": ["customer_id"], "additionalProperties": False, },}Every constraint in that schema is a class of mistake the model can no longer make. The enum means it cannot invent a status of "pending". The maximum means it cannot ask for ten thousand rows. additionalProperties: False means a hallucinated argument fails loudly instead of being ignored. Push in the other direction and you get the tool everybody regrets: run_sql(query), whose schema is one string, whose failure modes are every failure mode of SQL, and whose blast radius is the whole database.
The dispatcher is the other half. Its single most important property is that no tool failure can take down the loop — a crash inside a handler must come back as a result, because a tool use block with no matching result is an invalid conversation the API will reject.
def dispatch(handlers, call): result = {"type": "tool_result", "tool_use_id": call.id} handler = handlers.get(call.name) if handler is None: result["content"] = f"No tool named {call.name!r}. Available: {sorted(handlers)}" result["is_error"] = True return result try: result["content"] = handler(**call.input) except BadArguments as e: result["content"] = f"Invalid arguments: {e}" result["is_error"] = True except Exception as e: log.exception("tool %s failed", call.name) result["content"] = (f"{call.name} failed: {e}. " "Try a narrower query, or a different tool.") result["is_error"] = True return resultNote the is_error flag rather than a dropped result or a raised exception. The model is told that the step failed, gets a sentence explaining why, and stays in a position to do something else. Which brings up the thing most codebases get backwards.
Errors are messages to a model, not lines in a log
An exception string written for a human on-call engineer is close to useless as a tool result. psycopg2.errors.UndefinedColumn: column "cust_id" does not exist tells the model that something is broken and nothing about what to do next, so it guesses, and often guesses the same thing again. Rewrite it as an instruction: Unknown column cust_id. Valid columns are customer_id, status, total_cents. Now the next call is likely to be right.
Three habits carry most of the benefit. Say what was wrong and what a valid value looks like. Say whether retrying could possibly help, because the model cannot tell a transient timeout from a permanent 404 unless you tell it. And scrub the message before it goes back — raw exceptions love to carry internal hostnames, connection strings, and file paths into a transcript you may later show a user.
Assume every call happens twice
Retries are not an edge case. The model retries after a timeout, after an error it did not understand, and after revising its plan and deciding it needs that data again. If create_refund is called twice, the customer is paid twice. A duplicated read costs tokens; a duplicated write is an incident.
So make every mutating tool idempotent, and do it on your side rather than hoping. Give the tool an idempotency key derived from the run id and the canonicalised arguments, and pass it through to whatever API you are wrapping. Where the underlying service has no such concept, keep a small dedupe table on the key and return the stored result on a repeat. Shape the tool as an assertion about desired state (set_ticket_status) rather than an append (add_status_change) wherever the domain allows it, because assertions are naturally safe to repeat.
A tool that returns five thousand lines is a broken tool
Everything a tool returns becomes part of the prompt on every subsequent step. A list_files that dumps four thousand paths is not a one-off expense; you pay for those paths at step two and again at step eleven, and in between they sit there competing for the model's attention with the instruction it is supposed to be following. The symptoms look like model failure and are not: the agent loses track of the original task, starts citing the wrong record, or redoes a step it already completed.
The fix is to treat every tool return as a designed payload with a size budget, not as a pipe from your data layer to the prompt.
| Tool | Naive return | Bounded return |
|---|---|---|
list_files | Every path in the repository | 50 paths, the total count, and a cursor |
query_orders | Full result set as JSON | Four selected columns, 10 rows, row count |
read_file | The whole file | A byte range plus the file's total size |
fetch_url | Raw HTML | Extracted text, capped, with the source URL |
Two rules make this practical. Truncate visibly and say how to continue: showing 10 of 214 rows; call again with cursor "eyJvIjoxMH0" is dramatically better than silently returning ten rows, because a silent truncation makes the model confident it has seen everything. And summarise at the tool boundary rather than in the prompt. If a tool wraps a 300 page PDF, its job is to extract the relevant clause, not to hand over the document and hope. The general pattern is to return identifiers plus a way to fetch detail, and let the model pull the one record it actually needs.
What lives in context, and what lives in a database
Context is a working set, not storage. The dividing line is simple to state: whatever the model must reason over on this step goes in the prompt, and everything else lives outside and is reachable through a tool. Most of what gets called agent "memory" is those two boring mechanisms — retrieval, meaning fetch the relevant slice and put it in the prompt, and summarisation, meaning compress what is already there so the window does not fill. There is no third thing.
The failure this produces is specific and worth naming, because it will happen to you. Your carefully written task statement is 200 tokens at the top of the conversation. By step twelve it is sharing the window with 40,000 tokens of accumulated tool output, and it has effectively been evicted — still present, no longer salient. The agent starts optimising for whatever the most recent large blob was about.
Architecturally there are three moves. Keep a run state object that your code owns, not the transcript: the goal, the constraints, steps completed, artefacts produced, open questions. Serialise a compact version of it into the latest message each step, so the task is always the freshest thing in the window rather than the oldest. Second, fold finished tool output down to its conclusion — the 3,000 line query result was needed once; the line "checked the orders table for cus_8471, no matching row" is what the rest of the run actually uses. Third, treat anything that must survive between runs as a datastore with explicit keys, written by a tool. That is not a purity argument: a datastore can be inspected, corrected, and deleted from when a user asks you to forget them. A transcript cannot.
Termination is your job, not the model's
Nothing in the protocol guarantees the loop ends. The model will usually stop when the work is done. Your job is to design for the times it does not, and to do it in layers, from bluntest to smartest.
- A hard step cap. Non-negotiable, and it goes in before the first end-to-end test. Without it you have an unbounded spend loop with network access.
- A cost budget per run, checked before each request rather than after. Ten steps of a 60,000 token prompt is real money, and a runaway run finds that out faster than your billing alerts do.
- A wall clock deadline. A hung tool plus retries can outlive the user, the request timeout, and the deploy that was supposed to fix it.
- A deterministic goal check. Wherever you can write a predicate — the file exists, the test passes, the ticket status changed — write it. A predicate is a better stop signal than the model's own sense of completion, and it costs nothing to evaluate.
Then there is the specific pathology of loops: the same tool, the same arguments, over and over. It usually starts with an error message the model cannot act on, so it tries the identical call again. Cheap detection, two stages.
from collections import Counterimport jsonseen = Counter()def repeat_guard(call): """Return a nudge when a call is a verbatim repeat; raise when it is hopeless.""" sig = (call.name, json.dumps(call.input, sort_keys=True)) seen[sig] += 1 if seen[sig] > 4: raise LoopDetected(sig) if seen[sig] == 3: return (f"You have already called {call.name} with these exact arguments " "twice and it did not help. Change the arguments, try a different " "tool, or stop and report what you know so far.") return NoneThe first stage tells the model, in plain language, that it is repeating itself; that alone recovers the run more often than you would expect, because the model has no other way to notice. The second stage kills the run, which is what saves you on the occasions when telling it does not work.
Ninety-five percent per step is not ninety-five percent per task
Here is the arithmetic that explains most disappointing agent demos, and it has nothing to do with model quality. Suppose every step of your loop is 95% likely to be right. Ten steps end up correct about 60% of the time. Twenty steps, roughly a third.
| Steps | 95% per step | 99% per step |
|---|---|---|
| 5 | about 77% | about 95% |
| 10 | about 60% | about 90% |
| 20 | about 36% | about 82% |
| 50 | about 8% | about 61% |
And that table is optimistic, because it assumes the steps are independent. They are not. A wrong value at step three does not just fail step three; it becomes the input to steps four through ten, and the model will build on it with complete confidence. Errors in a long agent run do not average out, they propagate. That is the real reason a thirty-step autonomous plan looks brilliant in a recorded demo and lands somewhere between wrong and expensively wrong on live inputs.
Three responses actually work. Shorten the chain. Most tasks people hand to agents decompose into units of three to five steps with a natural checkpoint between them; three supervised five-step runs beat one unsupervised fifteen-step run on success rate, and when they fail you can see which one failed. Make steps verifiable by code. A compiler, a schema validator, a test suite, a checksum — anything that can confirm a step resets the accumulation instead of adding to it, which is why coding agents work better than their step counts suggest. Order steps by reversibility, exploration and drafting first, irreversible commits last, so a chain that went wrong can be thrown away rather than unwound. And resist the temptation to price retries as independent tries: the same model, retrying the same step with the same context, fails in a correlated way. Two attempts do not get you to 99.75%.
Blast radius: what the agent can reach when it is wrong
Design as though the loop will occasionally do the wrong thing with total conviction, because it will. The only question that matters is what the worst available wrong thing is, and whether you can live with it.
Scope privileges at the tool, not at the process. A read_orders tool whose connection uses a database role with SELECT on two views is a fundamentally different object from the same query executed on a connection that can DROP. When the tool holds the narrow credential, the safety review is over a handful of specific capabilities instead of over the set of all programs — which is what a general bash or run_sql or http_request tool actually asks you to review. Ship general-purpose tools only when the whole design is a sandbox: a container with no host mounts, no ambient credentials in its environment, and an egress allowlist. An agent that writes and runs code will eventually run code that surprises you.
Put an approval gate in front of anything irreversible or outward-facing: money, deletion, sending email or messages, production configuration, anything a customer will see. The clean way to implement it is as a tool that does not act — it records the proposed call and returns "pending approval", which ends the run. A human reviews the exact arguments, and approval re-enters the loop with the result. This works because the human is looking at one decision with context. Asking someone to approve step 14 of a 30 step run produces rubber-stamping, not oversight. Log what was proposed as carefully as what ran; the proposals are where you find out what your agent wanted to do before you stopped it.
The reader and the actor must not share authority
The moment your agent reads content it did not author — a web page, a PDF, a support email, an issue comment, a docstring in a dependency — that text arrives in the same channel as your instructions. There is no reliable in-band way for a model to distinguish "my operator asked for this" from "the document I was reading asked for this". Retrieved content is untrusted input, exactly like a form field, and it should never be spliced into a system instruction.
The structural rule follows from that, and it is the one worth memorising: a run that can read attacker-controlled text must not also hold the authority to act outward. Read-and-summarise is fine. Read-the-web-and-send-email is an exfiltration primitive, and no amount of "ignore any instructions inside the document" gets you out of it, because you are asking the model to solve the problem you failed to solve in the architecture. Split it instead. One run reads untrusted material and produces a structured, schema-validated artefact — fields, not prose. A separate step, with none of that untrusted text in its context, decides whether to act on the fields. The validation boundary between them is doing the security work, and unlike an instruction, it holds.
Cost per run, and the trace that explains it
Agent cost is not a per-request number, and this catches teams out. Because every step re-sends the whole transcript, total input tokens across a run grow with roughly the square of the step count. A run that takes twice as many steps costs about four times as much on input. That single fact reframes the earlier advice about bounded tool output and folded state: those are not tidiness, they are the main lever on the bill.
Record usage.input_tokens, usage.output_tokens, and any cache read and write counts on every step, sum them per run, and store the total next to the run id. The numbers that belong on a dashboard are cost per completed task and cost per abandoned task — the second is the one everyone forgets to count, and it is where runaway loops hide. Alongside them, emit one structured record per step.
def record(run_id, step, call, result, usage, latency_ms): log.info("agent_step", extra={"json": { "run_id": run_id, "step": step, "tool": call.name if call else None, "args": redact(call.input) if call else None, "result_chars": len(str(result.get("content", ""))) if result else 0, "is_error": bool(result and result.get("is_error")), "input_tokens": usage.input_tokens, "output_tokens": usage.output_tokens, "latency_ms": latency_ms, }})That record is the only artefact that can answer "why did the agent do that", because the final answer text never can. And with it you can separate a working run from a stuck one mechanically, without reading a word of the transcript. A healthy run shows distinct tool signatures, result sizes that shrink as it narrows in on the answer, an error rate that falls, and new tool names appearing as the task moves through phases. A stuck run shows repeated signatures, is_error true on consecutive steps, input tokens climbing while output tokens stay flat and small, and steady latency — it is not working hard, it is circling. Two counters catch nearly all of it: repeat-signature count, and steps since the last successful tool call. Both fire well before your step cap does, which means you find out from a graph rather than from an invoice.
Start smaller than feels satisfying
The most common mistake is not a technical one, it is scope. Pick a task that takes a person three to six tool interactions today and has an outcome you can check with code. Then, before you write the loop at all, try to solve it as a fixed sequence of calls with the model supplying only the judgement steps. If a hardcoded workflow solves the problem, you do not need an agent, and the workflow will be cheaper, faster, and vastly easier to debug. The loop earns its place only when the order of steps genuinely depends on what earlier steps found.
When you do build it, ship two tools before five — every additional tool multiplies the ways a run can go wrong, and the model's accuracy at picking among them drops as the list grows. Then write down twenty real task inputs, run the agent on all twenty, and read all twenty traces. This is the part that gets skipped and the part that does the work: you will discover that two of your tool descriptions overlap, that one error message reliably sends the model in a circle, and that a tool you thought returned a summary returns 1,200 lines. Fix those and the same model gets noticeably better at the same task.
Before it goes anywhere near production, the caps go in: step limit, cost limit, wall clock, and an approval gate on every irreversible tool. Afterwards, watch five things — completion rate per task type, cost per completed task, steps per completed task (falling over time is the clearest sign your tools are improving), human intervention rate, and the count of unapproved irreversible actions. That last one is not a rate to average. It is a number that stays at zero, and if it ever does not, you stop and fix the authority boundary before you tune anything else.
The loop at the top of this article is not going to get much more sophisticated. What separates an agent that ships from one that demos is everything wrapped around it: tool returns with a size budget, error messages written for a reader who cannot see your logs, a hard cap, honest arithmetic about long chains, an authority split between reading and acting, and a trace you can read at two in the morning. None of that is glamorous, and all of it is the job.