Course Content
Agents & Tools Interview Prep
6 sections · 40 lessons
How does an agent loop function, and what conditions determine when it stops?
What you need to know
One turn of the loop
- Call the model — send the messages so far plus the
toolslist. - Append the reply — add the full assistant content (text and
tool_useblocks) to the history. - Check why it stopped — read
stop_reason. - Run the tools — for each
tool_useblock, execute your own function. - Send results back — add one user message holding a
tool_resultblock per call, then go to step 1.
Here is the loop with the Anthropic Messages API. The model name and client live in one small module so no other file knows the provider:
1# llm.py — the only file that knows which provider and model we use2import os3import anthropic45MODEL = os.environ.get("AGENT_MODEL", "claude-opus-5")6_client = anthropic.Anthropic()78def call(messages, tools, system="", model=None, **kw):9 return _client.messages.create(10 model=model or MODEL, max_tokens=4096, system=system,11 tools=tools, messages=messages, **kw,12 )1import llm23def run_agent(task, tools, execute, max_steps=12):4 messages = [{"role": "user", "content": task}]5 for _ in range(max_steps):6 resp = llm.call(messages, tools)7 messages.append({"role": "assistant", "content": resp.content})8 if resp.stop_reason == "end_turn":9 return "".join(b.text for b in resp.content if b.type == "text")10 if resp.stop_reason != "tool_use": # max_tokens, refusal, ...11 raise RuntimeError(f"stopped early: {resp.stop_reason}")12 results = [execute(b) for b in resp.content if b.type == "tool_use"]13 messages.append({"role": "user", "content": results}) # all results, one message14 raise RuntimeError("step limit reached")execute returns a dict like {"type": "tool_result", "tool_use_id": b.id, "content": "..."}. Two details matter: append the whole resp.content (not just the text), and send all results of one turn in one user message.
The stop reasons you must handle
stop_reason | Meaning | What the loop does |
|---|---|---|
end_turn | Model finished its answer | Return the text |
tool_use | Model wants tools run | Run them, continue |
max_tokens | Reply was cut off | Raise the limit or fail; never run a half-written tool call |
stop_sequence | Hit a custom stop string | Usually treat as done |
pause_turn | A server-side tool loop paused | Send the conversation back to resume |
refusal | Model declined for safety reasons | Stop; do not run any tools from that turn |
OpenAI has the same idea: in Chat Completions, finish_reason is "tool_calls"; in the Responses API, the output contains function_call items.
Harness-driven stops
The model deciding "I'm done" is necessary but not sufficient. Add these yourself:
- Step cap — e.g. 12 model calls.
- Budget — total tokens or rupees/dollars, from each response's
usage. - Deadline — wall-clock time, e.g. 60 seconds for a chat user.
- Repetition detector — the same tool with the same arguments three times means stuck.
- Error streak — three tool failures in a row ends the run with a clear message.
When a guard fires, return a useful partial answer ("I found the flights but could not confirm hotel prices") rather than a stack trace.
A real-life example
A GitHub triage bot labels new issues. For issue #4812, the model calls search_issues("login timeout"), gets 3 similar issues, calls get_issue(4790), sees it is the same bug, and calls add_label and add_comment ("Duplicate of #4790"). Then it replies with a summary and stop_reason is end_turn. Four tool calls, done.
A week later, a repository's search API starts returning an empty list because of an expired token. The model keeps calling search_issues with slightly different words — 40 calls in one run, until the 12-step cap fires. The cap turned an unbounded bill into a bounded one; a repetition detector, added afterwards, now stops such runs after 3 near-identical calls, and the empty-result bug is fixed by making the tool return an error ("GitHub token expired") instead of [].
Follow-up questions to expect
- "Why append the whole assistant content, not just text?" — Each
tool_resultmust refer to atool_useid in the previous assistant message. Drop thetool_useblocks and the next request fails. - "Should the agent get a
finishtool?" — Useful when you need a structured final answer: the model callssubmit_answer(...)with a schema, and the loop ends there. - "Do I have to write this loop myself?" — No. SDK helpers such as Anthropic's tool runner, or frameworks like LangGraph, run it for you. You still set the limits.