Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your agent calls a search tool, doesn't like the results, calls it again with the same query, and loops forever. How do you detect and break agent loops?


The run that used to loopsearch 'mobilecovers': 0 resultssearch 'mobilecovers': repeattool returnsrepeat_call hintsearch 'phonecases': 12 resultsanswer tothe usernullsame hash,not executedmodel changes course
Returning the repeat as a readable tool result, not an exception, is what lets the model recover.

What you need to know

Why does the model repeat itself? It sees a result it cannot use, such as an empty list with no explanation, and its best guess is "try again". Nothing in the conversation tells it that trying again will give the same answer.

Level 1: hard bounds

These guarantee the run ends, whatever the model does.

BoundTypical valueWhen it trips
Step cap8 to 12 steps for most tasksToo many model turns
Wall-clock deadline30 to 60 seconds for chatSlow tools or long reasoning
Token budgetSet per task typeExpensive runs

Whichever trips first ends the run with the best partial answer and a clear message, not a hang.

Level 2: repetition detection

Python
def check_repeat(name: str, args: dict, seen: set) -> dict | None:    sig = hashlib.sha256(f"{name}:{json.dumps(args, sort_keys=True)}".encode()).hexdigest()    if sig in seen:        return {"error": "repeat_call",                "message": "You already called this tool with these exact arguments. "                           "Use the earlier result, change the query, or answer with what you have."}    seen.add(sig)    return None

Return this as a tool result, not an exception. The model reads it and changes course. Also watch near-repeats: if three steps in a row use the same tool with slightly reworded queries, insert a reflection step asking the model to summarise what it has learned and decide whether to answer.

Level 3: fix the cause

Look at the tool output just before each loop starts. Most of the time it is one of:

  • An empty result with no explanation: [].
  • A bare error string the model does not know how to act on: "500".
  • A result too large or noisy to read, so the model tries again hoping for better.

Make results self-describing:

JSON
{"results": [], "hint": "No matches for 'refund policy 2019'. Try broader terms or remove the year."}

With a hint, the model changes its query or answers "I couldn't find that", and loop frequency drops sharply.

  1. Bound — step cap, deadline and budget on every run.
  2. Detect — hash tool name and arguments; return a repeat message.
  3. Diagnose — read traces before loops; find the tool result that caused them.
  4. Fix the tool — self-describing empty results and errors.

Metrics

Loop-abort rate, the distribution of steps per task (a second bump at the step cap is the tell), and cost per task. Alert on abort rate; a rise is often the first sign that a tool broke.

A real-life example

Scenario (illustrative numbers). An e-commerce company's shopping assistant uses a product-search tool. After a catalogue migration, searches with old category names return []. The agent retries the same search, and 6% of sessions run until a 120-second timeout, each costing about 25 model calls.

The team adds a 10-step cap, a 45-second deadline and repeat detection. The same day, runaway sessions stop. The traces show the empty results come from 40 renamed categories, so the search tool now returns a hint with the new category names. Average steps per task fall from 5.4 to 2.9, and the loop-abort rate settles at 0.3%.

Follow-up questions to expect

  • "Why not just lower the step cap?" — A cap guarantees termination but ends good long tasks too; detection and better tool results fix the behaviour, and the cap stays as a backstop.
  • "How do you handle loops across two tools, A then B then A?" — Track the sequence of recent signatures and detect cycles, not only exact repeats.
  • "What does the user see when a run is stopped?" — The best partial answer, plus an honest line such as "I couldn't find X; here's what I did find", or a hand-off to a human.