Applied AI Engineering: From Prompt to Production

Course Content

Applied AI Engineering: From Prompt to Production

9 sections · 29 lessons

The agent loop and planning


Look again at Priya's message from the first lesson: she is back from parental leave, wants to know about work-from-home days, and has a VPN problem. One tool round cannot handle that. PolicyPal must search the return-to-work policy, possibly search the remote-access standard, answer, and propose a ticket. The model cannot know in advance exactly which steps it will need, because the second step depends on what the first one returns.

That is what an agent loop is for: call the model, run any tools it asks for, send back the results, and repeat until it gives a final answer. It is a short piece of code. The hard part is everything that keeps it from going wrong.

In early testing, one run searched the policies eleven times with slightly reworded queries, used 38,000 tokens, took 41 seconds and still gave a vague answer. Another run raised a ticket the user never asked for, because a retrieved paragraph said "raise a ticket for any VPN issue". This lesson builds the loop and the guards against both.

One turn of PolicyPal's bounded loopcall the modelfinalanswer? stopsideeffect? stop, askrun tools,block repeatscheck steps,tokens, timeconfirmation card6 steps, 20ktokens, 20 sOne unguarded run made 11 searches, used 38,000 tokens and took 41 seconds.
The loop is a few lines; the guards around it lifted success from 78% to 92% with no unconfirmed tickets.

A bounded loop

An agent loop needs three budgets, each checked on every turn: a maximum number of steps, a maximum number of tokens, and a deadline. Without them, a confused model can loop until someone notices the bill.

Python
# policypal/agent.pyimport timefrom dataclasses import dataclassfrom policypal.tool_defs import TOOLSfrom policypal.tools import run_tool, tool_resultSIDE_EFFECTS = {"create_it_ticket"}GIVE_UP = "I couldn't finish this reliably. I've passed your question to the helpdesk."@dataclassclass AgentResult:    text: str    pending: dict | None = None          # an action waiting for the user's confirmation    steps: int = 0    tokens: int = 0def run_agent(llm, system: str, messages: list[dict], user, *, max_steps: int = 6,              max_tokens: int = 20_000, deadline_s: float = 20.0) -> AgentResult:    start, tokens, seen = time.monotonic(), 0, set()    for step in range(1, max_steps + 1):        reply = llm.complete(system, messages, tools=TOOLS, max_tokens=800)        tokens += reply.input_tokens + reply.output_tokens        if reply.stop_reason != "tool_use":            return AgentResult(reply.text, steps=step, tokens=tokens)        for call in reply.tool_calls:            if call["name"] in SIDE_EFFECTS:           # stop and ask the human                return AgentResult(reply.text, pending=call, steps=step, tokens=tokens)        results = []        for call in reply.tool_calls:            key = (call["name"], repr(sorted(call["input"].items())))            if key in seen:                results.append(tool_result(call["id"], {"error": "Same call already made; "                               "use the earlier result."}, error=True))            else:                seen.add(key)                results.append(run_tool(user, call))        messages = messages + [{"role": "assistant", "content": reply.content},                               {"role": "user", "content": results}]        if tokens > max_tokens or time.monotonic() - start > deadline_s:            break    return AgentResult(GIVE_UP, steps=step, tokens=tokens)

Each turn does the same thing: call the model, stop if it gave a final answer, otherwise run every tool call it made and send all results back together. The seen set catches exact repeats: the second identical call gets an error result telling the model to use what it already has, which usually breaks the cycle. When a budget runs out, the loop does not return a half-finished guess. It returns an honest message and, in production, opens a helpdesk case with the transcript.

Six steps, 20,000 tokens and 20 seconds are PolicyPal's numbers, set from measurement: 95% of successful runs on the 40 test tasks finished within 4 steps and 11,000 tokens. Budgets should sit comfortably above normal behaviour, so they only stop abnormal runs.

Confirmation before side effects

The loop stops the moment the model asks for a tool in SIDE_EFFECTS. It returns the model's text and the proposed call as pending, without running it. The web page shows the pending ticket as a card with the exact category, summary and urgency, and two buttons: Confirm and Cancel.

When the user clicks Confirm, the application runs run_tool(user, pending) directly, with exactly the arguments shown on the card. It does not send "yes" back to the model and ask it to try again. That matters: if the model re-generated the call after "yes", it could produce slightly different arguments from the ones the user approved. The user confirms a specific action, and that specific action runs. The result, or "cancelled by user", is then added to the conversation as the tool_result for that call, so the model knows what happened on the next turn.

This one rule also defuses the injected-paragraph problem from the opening. Even if a retrieved document persuades the model to propose a ticket, nothing happens until a person reads the card and clicks. Section 9 adds more layers, but this is the one that makes a write action safe to offer at all.

Planning: how much do you need?

For long tasks, such as an agent that researches and writes a report, an explicit plan-then-execute step helps: the model first writes a list of steps, then works through it, and you can check the plan before any tool runs. PolicyPal's tasks are short, usually two to four steps, and an explicit planning call added a second of latency without improving success on the test tasks.

What did help was one sentence in the system prompt: "Decide which tools you need. Call tools that do not depend on each other in the same turn." Models can return several tool calls in one reply, called parallel tool calls, and the loop already runs all of them before the next model call. For Priya's message, both policy searches now happen in one turn. Across the 40 tasks, the median number of steps fell from 3.9 to 2.6, and median time from 11 seconds to 7.

Failure modes and their guards

FailureWhat you seeGuard
RepetitionThe same call again and againThe seen set returns an error for exact repeats
Flailing searchMany reworded searches, no answerAt most 3 searches per turn; after that the tool returns "no further results"
Runaway costA long run with growing contextToken and time budgets; trimmed tool results
Invented argumentsA ticket category the user never saidEnums in schemas, and the confirmation card
Acting on retrieved textA tool call that a document "asked for"Side effects need a human click
Stopping too earlyAnswers after one search, missing half the questionTest tasks with two intents; prompt says to cover every part

The flailing-search guard deserves a note, because it fixed the 38,000-token run. The model was not stupid; the answer simply was not in the policies, and nothing told it to stop looking. After three searches, the tool now returns "No further results. If the answer is not in what you have, say it is not in the policy." The model does exactly that.

Check your understanding

0 of 3 answered

1.When the user clicks Confirm on a pending ticket, why does PolicyPal run the stored call directly instead of sending "yes" to the model?

2.A run searches the policies 11 times with reworded queries and never answers. What is the most likely cause and fix?

3.Why are PolicyPal's budgets set at 6 steps and 20,000 tokens when 95% of successful runs use at most 4 steps and 11,000 tokens?