Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Scenario – 2: Cyclic Execution Loop
What you need to know
The scenario: an agent graph keeps cycling between the model node and the tool node, burning tokens until it times out or hits a limit.
Why agents loop
The standard agent is a cycle: the model decides, a tool runs, the result goes back to the model. The loop ends when the model stops asking for tools. It spins forever when that exit can never be reached — most often because a tool keeps failing, the failure is hidden from the model, and nothing in state changes between turns.
Three layers of termination
| Layer | What it does | Result for the user |
|---|---|---|
| Counter in state | Routes to a finish node after N iterations | A graceful partial answer |
recursion_limit in config | Hard stop from the runtime | An exception you catch and turn into the best answer so far |
| Progress detection | Stops when the same call repeats | Stops loops early, before the cap |
1from typing import Literal2from langgraph.errors import GraphRecursionError34def should_continue(state) -> Literal["tools", "finish"]:5 if state["iterations"] >= 5:6 return "finish" # graceful exit7 last = state["messages"][-1]8 if not last.tool_calls:9 return "finish"10 sig = [(c["name"], str(c["args"])) for c in last.tool_calls]11 if sig == state.get("last_call_sig"):12 return "finish" # same call again: no progress13 return "tools"1415try:16 result = graph.invoke(inputs, {"recursion_limit": 30, "configurable": {"thread_id": tid}})17except GraphRecursionError:18 result = graph.get_state({"configurable": {"thread_id": tid}}).values # best so farThe tool node should update iterations and last_call_sig in state, so the router can see them.
Fix the root cause
Loops are a symptom. Feed tool errors back into the message history as tool results, so the model sees "invoice API returned 404" and can change course. Cap retries per tool. Make the finish node honest: say what was tried and what is missing.
Watch it
Track steps per run (p95), the GraphRecursionError rate and the repeated-call rate. In a tracing tool such as LangSmith, a loop shows up as the same pair of spans repeating.
A real-life example
Scenario, numbers made up. A travel agent graph books trains by calling a seat-availability tool. When the tool's upstream returns a 503, the tool wrapper returns an empty list, and the model asks again, and again. Some runs take 40 steps and cost about 30 times a normal run.
The team adds the three layers: a cap of 5 iterations, recursion_limit of 30, and repeated-call detection. They change the tool to return "availability service unavailable, try later" instead of an empty list. The model now tells the user the service is down and offers to retry later. p95 steps per run falls from 14 to 4, and cost per conversation drops by about 40%.
Follow-up questions to expect
- "Isn't the recursion limit enough?" — It stops the spin, but as an exception after many wasted steps. The counter and progress check stop earlier and produce a proper answer.
- "What should the finish node say?" — What was tried, what failed, and what the user can do next; a partial answer with honesty beats a timeout.
- "How do you choose the cap?" — From traces: look at how many steps successful runs take, and set the cap a little above their p99.