Course Content
Agents & Tools Interview Prep
6 sections · 40 lessons
How do you detect and fix infinite loops in agent execution?
What you need to know
Three kinds of loop
| Kind | Looks like | Detector |
|---|---|---|
| Exact repeat | search_issues("login bug") × 5 | Count hashes of (tool, sorted-JSON args) |
| Cycle | get_order → get_customer → get_order → … | Repeating subsequence in the last N calls |
| Drift | New queries each time, nothing new learned | No progress for K steps |
A detector that nudges before it kills
1import json2from collections import Counter34class LoopGuard:5 def __init__(self, warn_at=2, stop_at=3):6 self.seen, self.warn_at, self.stop_at = Counter(), warn_at, stop_at78 def check(self, block):9 key = (block.name, json.dumps(block.input, sort_keys=True))10 self.seen[key] += 111 n = self.seen[key]12 if n >= self.stop_at:13 raise LoopDetected(block.name)14 if n >= self.warn_at:15 return {"type": "tool_result", "tool_use_id": block.id, "is_error": True,16 "content": f"You already called {block.name} with these exact arguments "17 f"and the result will not change. Try a different approach, "18 f"or answer with what you know and say what is missing."}19 return None # fine, run the tool normallyReturning the nudge as an is_error tool result usually breaks the loop and keeps a useful answer. If the model repeats anyway, the run stops cleanly.
Fix the causes, not just the symptom
Loops usually come from tools, not from the model being "dumb":
- Unhelpful empty results.
[]gives no clue why. Say "No issues match 'login bug' in repo X. Try broader words or a different label." - Errors with no way forward. "Error 400" makes the model retry the same thing. "
datemust be in the future" tells it what to change. - Overlapping tools. Two tools that each seem half-right invite ping-pong.
- No clear finish. If the prompt never says what "done" looks like, the model keeps searching for more.
A real-life example
A GitHub triage bot sometimes ran to its 12-step cap. The traces showed a pattern: it called search_issues with a query, got [], rephrased, got [], and so on — 11 searches in one run.
The root cause: the search tool silently returned [] when the query contained a colon, because the GitHub search syntax treated error: timeout as a qualifier. The bot could never succeed.
Fixes: the tool now escapes the query, and when it gets zero results it says "0 results for query X; searched open and closed issues in repo Y". The loop guard nudges on the second identical call and stops on the third. Runs hitting the cap fell from 6% to 0.4%, and average cost per issue fell by a third.
Follow-up questions to expect
- "Why not just lower the step cap?" — It limits damage but also blocks legitimate long runs. Detectors stop bad runs early while allowing good long ones.
- "How do you measure progress for drift?" — Count new unique entities or facts in tool results, or plan steps completed; no change for 3–4 steps means stuck.
- "Should the loop detector live in the prompt?" — No. The model cannot reliably count its own calls; the harness can.