Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your customer support agent enters an infinite tool-calling loop: search → summarize → search again forever. How do you detect and stop recursive agent failure loops in production?
What you need to know
Why agents loop
An agent loop is: the model picks a tool, reads the result, and picks again. It loops when nothing in the result tells it to change strategy. Three causes cover most cases:
- Unhelpful tool results. An empty list
[]gives no reason to try something different, so the model searches again with a synonym. - Hidden history. After twenty messages, the model's own earlier attempts are buried, and it forgets it already tried "refund status".
- An impossible goal. The prompt says "find the answer" and gives no allowed way to stop without one.
Detection signals, in layers
| Signal | What it catches | Cost |
|---|---|---|
| Step cap and token/time budget | Everything, eventually | Free; a backstop only |
Hash of (tool, normalised args) | Exact repeats and short cycles | Free |
| Semantic no-progress check | "Search again with a synonym" | One embedding per step |
| Cost per conversation alert | Loops you did not predict | Monitoring only |
1import hashlib, json23class LoopGuard:4 def __init__(self, max_steps=15, window=6):5 self.max_steps, self.window = max_steps, window6 self.calls, self.states = [], []78 def check(self, tool: str, args: dict, findings_vec) -> str | None:9 key = hashlib.sha1(f"{tool}:{json.dumps(args, sort_keys=True).lower()}".encode()).hexdigest()10 if key in self.calls[-self.window:]:11 return "repeat_call"12 self.calls.append(key)13 self.states.append(findings_vec)14 if len(self.calls) > self.max_steps:15 return "step_cap"16 if len(self.states) >= 4 and all(cosine(self.states[-1], s) > 0.97 for s in self.states[-4:-1]):17 return "no_progress" # findings stopped changing18 return Nonefindings_vec is an embedding of the notes gathered so far, and cosine is cosine similarity. If three steps in a row add nothing new, the state vector hardly moves and the guard stops the run. The 0.97 threshold is a starting value to tune on your own traces.
Fix the causes
- Make empty results useful — return "0 results for 'refund status'; the index covers orders from 2023 onward; try the order id." The model can act on that.
- Show what was tried — keep an
attempted_actionslist near the end of the context, where the model will read it. - Allow an honest ending — "I couldn't find X; here is what I checked" is a valid final answer.
- Always end with something — on any guard trip, go to a terminal step that writes a partial answer, and escalate the ticket to a human after repeated failures.
Watch p50 and p95 steps per run, cap-hit rate, repeat-call rate and cost per resolved ticket.
A real-life example
Scenario, numbers made up. A telecom support agent handles 40,000 chats a day. Tracing shows 3% of runs hit the 25-step cap, each costing about 18 times a normal run, and ending in a timeout message.
Most loops share one pattern: the knowledge-base search returns an empty list for plan names discontinued last year, and the agent keeps rephrasing. The team changes the tool to say "no results; plan names before 2024 are in the legacy-plans index", adds the attempted-actions list and the loop guard, and lets the agent end with "I couldn't find this plan; connecting you to an agent." Cap hits fall to 0.3%, and those runs now reach a human in under a minute instead of timing out.
Follow-up questions to expect
- "How do you pick the step cap?" — From traces of healthy runs: set it a little above their p99, so it only trips on real failures.
- "How do you tell a genuine long search from a loop?" — Progress. A long run that keeps adding new findings is fine; a short run whose findings stopped changing is not.
- "Does your framework handle this?" — Partly. LangGraph, for example, raises an error when a run passes its recursion limit, but that is only the step-cap backstop, not progress detection.