Agents & Tools Interview Prep

Course Content

Agents & Tools Interview Prep

6 sections · 40 lessons

How does an agent loop function, and what conditions determine when it stops?


One issue, four tool calls, one clean stopModel call withmessages and toolsstop_reason:tool_useHarness runssearch, get, labelAll tool_resultblocks in one messagestop_reason:end_turn —return textStep cap, budget, deadline and repeat detector sit outside this loop.
The model decides when it is done, but only the harness can guarantee the loop ends when it is not.

What you need to know

One turn of the loop

  1. Call the model — send the messages so far plus the tools list.
  2. Append the reply — add the full assistant content (text and tool_use blocks) to the history.
  3. Check why it stopped — read stop_reason.
  4. Run the tools — for each tool_use block, execute your own function.
  5. Send results back — add one user message holding a tool_result block per call, then go to step 1.

Here is the loop with the Anthropic Messages API. The model name and client live in one small module so no other file knows the provider:

Python
# llm.py — the only file that knows which provider and model we useimport osimport anthropicMODEL = os.environ.get("AGENT_MODEL", "claude-opus-5")_client = anthropic.Anthropic()def call(messages, tools, system="", model=None, **kw):    return _client.messages.create(        model=model or MODEL, max_tokens=4096, system=system,        tools=tools, messages=messages, **kw,    )
Python
import llmdef run_agent(task, tools, execute, max_steps=12):    messages = [{"role": "user", "content": task}]    for _ in range(max_steps):        resp = llm.call(messages, tools)        messages.append({"role": "assistant", "content": resp.content})        if resp.stop_reason == "end_turn":            return "".join(b.text for b in resp.content if b.type == "text")        if resp.stop_reason != "tool_use":          # max_tokens, refusal, ...            raise RuntimeError(f"stopped early: {resp.stop_reason}")        results = [execute(b) for b in resp.content if b.type == "tool_use"]        messages.append({"role": "user", "content": results})  # all results, one message    raise RuntimeError("step limit reached")

execute returns a dict like {"type": "tool_result", "tool_use_id": b.id, "content": "..."}. Two details matter: append the whole resp.content (not just the text), and send all results of one turn in one user message.

The stop reasons you must handle

stop_reasonMeaningWhat the loop does
end_turnModel finished its answerReturn the text
tool_useModel wants tools runRun them, continue
max_tokensReply was cut offRaise the limit or fail; never run a half-written tool call
stop_sequenceHit a custom stop stringUsually treat as done
pause_turnA server-side tool loop pausedSend the conversation back to resume
refusalModel declined for safety reasonsStop; do not run any tools from that turn

OpenAI has the same idea: in Chat Completions, finish_reason is "tool_calls"; in the Responses API, the output contains function_call items.

Harness-driven stops

The model deciding "I'm done" is necessary but not sufficient. Add these yourself:

  • Step cap — e.g. 12 model calls.
  • Budget — total tokens or rupees/dollars, from each response's usage.
  • Deadline — wall-clock time, e.g. 60 seconds for a chat user.
  • Repetition detector — the same tool with the same arguments three times means stuck.
  • Error streak — three tool failures in a row ends the run with a clear message.

When a guard fires, return a useful partial answer ("I found the flights but could not confirm hotel prices") rather than a stack trace.

A real-life example

A GitHub triage bot labels new issues. For issue #4812, the model calls search_issues("login timeout"), gets 3 similar issues, calls get_issue(4790), sees it is the same bug, and calls add_label and add_comment ("Duplicate of #4790"). Then it replies with a summary and stop_reason is end_turn. Four tool calls, done.

A week later, a repository's search API starts returning an empty list because of an expired token. The model keeps calling search_issues with slightly different words — 40 calls in one run, until the 12-step cap fires. The cap turned an unbounded bill into a bounded one; a repetition detector, added afterwards, now stops such runs after 3 near-identical calls, and the empty-result bug is fixed by making the tool return an error ("GitHub token expired") instead of [].

Follow-up questions to expect

  • "Why append the whole assistant content, not just text?" — Each tool_result must refer to a tool_use id in the previous assistant message. Drop the tool_use blocks and the next request fails.
  • "Should the agent get a finish tool?" — Useful when you need a structured final answer: the model calls submit_answer(...) with a schema, and the loop ends there.
  • "Do I have to write this loop myself?" — No. SDK helpers such as Anthropic's tool runner, or frameworks like LangGraph, run it for you. You still set the limits.