Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Scenario – 2: Cyclic Execution Loop


Three exits for one agent loopCounter in state: finish after 5 turnsRepeated identical tool call: stoprecursion_limit: hard runtime stopTool errors fed back to the model
The caps stop the spin, but the loop only ends properly once the model can see that the tool failed.

What you need to know

The scenario: an agent graph keeps cycling between the model node and the tool node, burning tokens until it times out or hits a limit.

Why agents loop

The standard agent is a cycle: the model decides, a tool runs, the result goes back to the model. The loop ends when the model stops asking for tools. It spins forever when that exit can never be reached — most often because a tool keeps failing, the failure is hidden from the model, and nothing in state changes between turns.

Three layers of termination

LayerWhat it doesResult for the user
Counter in stateRoutes to a finish node after N iterationsA graceful partial answer
recursion_limit in configHard stop from the runtimeAn exception you catch and turn into the best answer so far
Progress detectionStops when the same call repeatsStops loops early, before the cap
Python
from typing import Literalfrom langgraph.errors import GraphRecursionErrordef should_continue(state) -> Literal["tools", "finish"]:    if state["iterations"] >= 5:        return "finish"                                  # graceful exit    last = state["messages"][-1]    if not last.tool_calls:        return "finish"    sig = [(c["name"], str(c["args"])) for c in last.tool_calls]    if sig == state.get("last_call_sig"):        return "finish"                                  # same call again: no progress    return "tools"try:    result = graph.invoke(inputs, {"recursion_limit": 30, "configurable": {"thread_id": tid}})except GraphRecursionError:    result = graph.get_state({"configurable": {"thread_id": tid}}).values   # best so far

The tool node should update iterations and last_call_sig in state, so the router can see them.

Fix the root cause

Loops are a symptom. Feed tool errors back into the message history as tool results, so the model sees "invoice API returned 404" and can change course. Cap retries per tool. Make the finish node honest: say what was tried and what is missing.

Watch it

Track steps per run (p95), the GraphRecursionError rate and the repeated-call rate. In a tracing tool such as LangSmith, a loop shows up as the same pair of spans repeating.

A real-life example

Scenario, numbers made up. A travel agent graph books trains by calling a seat-availability tool. When the tool's upstream returns a 503, the tool wrapper returns an empty list, and the model asks again, and again. Some runs take 40 steps and cost about 30 times a normal run.

The team adds the three layers: a cap of 5 iterations, recursion_limit of 30, and repeated-call detection. They change the tool to return "availability service unavailable, try later" instead of an empty list. The model now tells the user the service is down and offers to retry later. p95 steps per run falls from 14 to 4, and cost per conversation drops by about 40%.

Follow-up questions to expect

  • "Isn't the recursion limit enough?" — It stops the spin, but as an exception after many wasted steps. The counter and progress check stop earlier and produce a proper answer.
  • "What should the finish node say?" — What was tried, what failed, and what the user can do next; a partial answer with honesty beats a timeout.
  • "How do you choose the cap?" — From traces: look at how many steps successful runs take, and set the cap a little above their p99.