Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Scenario – 7: Agent Infinite Loop Problem
Scenario: a LangChain agent keeps calling tools without converging, and either runs until the timeout or burns thousands of tokens per request. How do you fix it?
What you need to know
An agent loops when it keeps choosing another step. Something in code has to be able to say "stop", because the model will not reliably say it when it is confused.
Bounds, in current and legacy APIs
1from langchain.agents import create_agent2from langchain.agents.middleware import ModelCallLimitMiddleware, ToolCallLimitMiddleware34agent = create_agent(5 model=llm, tools=tools,6 middleware=[ModelCallLimitMiddleware(run_limit=10, exit_behavior="end"),7 ToolCallLimitMiddleware(tool_name="search", run_limit=4)])89result = agent.invoke({"messages": [("user", question)]},10 config={"recursion_limit": 25}) # hard cap on graph stepsModelCallLimitMiddleware ends the run after 10 model calls; ToolCallLimitMiddleware limits one tool to 4 calls per run. recursion_limit is LangGraph's backstop on total steps. Add a wall-clock timeout around the call as well.
For older code on the legacy AgentExecutor (now in langchain-classic), the same bounds are max_iterations=8, max_execution_time=45 and early_stopping_method. They exist but are often left at defaults. New work should use create_agent.
Repeat detection
Keep a set of (tool name, sorted arguments) signatures per run. When one repeats, return a tool message instead of running the tool:
You already called search with {"query": "refund policy"} and got 0 results.Use that result, try a different query, or answer with what you have.Returning it as a tool observation, not an exception, is what lets the model recover. This can live in a custom middleware or inside the tool wrapper.
Fix the cause
- Find loops — in LangSmith, filter runs by step count near the limit.
- Read the observations — look at the tool result right before repetition starts.
- Classify — usually an empty list, a bare exception string, or a huge unreadable payload.
- Fix the tool — return
{"results": [], "hint": "no match; try a broader query"}or a clear, actionable error.
Most of the time, the model loops because the tool gave it nothing to reason about.
Metrics
Distribution of steps per run (a bump at the limit is the tell), loop-abort rate, and cost per run. Alert on the abort rate; it often rises when a downstream API changes.
A real-life example
Scenario (illustrative numbers). A property portal's agent answers questions like "3BHK in Whitefield under ₹1.2 crore". It has no limits set. After the listings API changes its locality names, searches for "Whitefield" return an empty list, and 5% of runs call search_listings 20 or more times, costing about ₹40 each instead of ₹3.
The team moves to create_agent with a 10-call model limit, a 4-call limit on the search tool, a 30-second deadline and repeat detection. Runaway runs stop the same day. The traces reveal the empty results, so the tool now returns a hint with the closest locality names. Loop aborts fall to 0.2%, and average cost per run returns to about ₹3.
Follow-up questions to expect
- "What should the user see when a limit trips?" — The best partial answer plus a clear statement of what couldn't be done, or a hand-off to a human; never a raw error.
- "How do you choose the limits?" — From the step distribution of successful runs: set the cap a little above the 99th percentile.
- "Can the model itself detect it's stuck?" — A reflection step helps, but it is advice to the model; the code-level limits are the guarantee.