Course Content
Live Coding Interview Prep
7 sections · 50 lessons
Write an agent loop that decides when to stop execution.
What you need to know
"Stop when the model says it is finished" is only the happy path. Agents fail by not stopping: they loop on the same search, wander, or burn the budget on a task they cannot do. A robust loop has two kinds of stop:
- Success — the model emits a final answer.
- Budgets and safety valves — steps, seconds, tokens, money, and "no progress". Each is a separate check because each catches a different failure.
Fingerprinting an action means turning it into a short, comparable string. json.dumps(action, sort_keys=True) gives the same text for the same action no matter the key order; hashing it keeps the stored value small.
time.monotonic() is a clock that only moves forward. time.time() follows the system clock, which NTP can adjust mid-run.
1import hashlib, json, time2from collections import deque3from collections.abc import Callable4from dataclasses import dataclass, field56@dataclass7class Budget:8 max_steps: int = 109 max_seconds: float = 60.010 max_tokens: int = 50_00011 max_usd: float = 0.5012 repeat_limit: int = 31314@dataclass15class StopController:16 """Tracks spend and recent actions; should_stop() returns a reason or None."""17 budget: Budget = field(default_factory=Budget)18 steps: int = 019 tokens: int = 020 usd: float = 0.021 started: float = field(default_factory=time.monotonic)22 recent: deque = field(default_factory=lambda: deque(maxlen=10))2324 def record(self, action: dict, tokens: int = 0, usd: float = 0.0) -> None:25 self.steps, self.tokens, self.usd = self.steps + 1, self.tokens + tokens, self.usd + usd26 key = json.dumps(action, sort_keys=True, default=str)27 self.recent.append(hashlib.sha256(key.encode()).hexdigest()[:16])2829 def should_stop(self) -> str | None:30 b = self.budget31 if self.steps >= b.max_steps:32 return "step limit"33 if time.monotonic() - self.started > b.max_seconds:34 return "time limit"35 if self.tokens >= b.max_tokens:36 return "token budget"37 if self.usd >= b.max_usd:38 return "cost budget"39 last = list(self.recent)[-b.repeat_limit:]40 if len(last) == b.repeat_limit and len(set(last)) == 1:41 return "no progress: same action repeated"42 return None4344def agent_loop(question: str, policy: Callable, execute: Callable,45 controller: StopController) -> dict:46 """policy(transcript) -> action dict; execute(action) -> observation dict."""47 transcript: list = [("question", question)]48 while (reason := controller.should_stop()) is None:49 action = policy(transcript)50 if action.get("tool") == "final_answer":51 return {"answer": action.get("answer"), "stopped": "completed", "steps": controller.steps}52 observation = execute(action)53 controller.record(action, tokens=observation.get("tokens", 0), usd=observation.get("usd", 0.0))54 transcript.append((action, observation))55 return {"answer": None, "stopped": reason, "steps": controller.steps}The tricky parts:
- The check runs before each step, so a budget that is already spent never triggers one more expensive call.
- The repeat check looks at the last
repeat_limitfingerprints only. Callingsearch("refund")three times in a row is stuck; calling it at steps 1 and 7 may be legitimate. sort_keys=Truemakes{"a": 1, "b": 2}and{"b": 2, "a": 1}the same fingerprint.- The walrus
:=reads the reason and tests it in one line, so the final return can report which limit fired.
Complexity: record hashes one action, O(size of the action). should_stop is O(repeat_limit), which is constant. The deque(maxlen=10) keeps memory O(1) however long the run is.
A real-life example
A stuck policy that keeps searching for the same thing:
1stuck = lambda transcript: {"tool": "search", "args": {"query": "refund policy"}}2execute = lambda action: {"result": "no results", "tokens": 1200}3print(agent_loop("What is the refund policy?", stuck, execute, StopController()))4# {'answer': None, 'stopped': 'no progress: same action repeated', 'steps': 3}56tight = StopController(budget=Budget(max_tokens=3000))7varied = lambda t: {"tool": "search", "args": {"query": f"refund try {len(t)}"}}8print(agent_loop("What is the refund policy?", varied, execute, tight))9# {'answer': None, 'stopped': 'token budget', 'steps': 3}First run, step by step:
| before step | steps | tokens | last 3 fingerprints | should_stop |
|---|---|---|---|---|
| 1 | 0 | 0 | – | None → run |
| 2 | 1 | 1,200 | A | None → run |
| 3 | 2 | 2,400 | A, A | None → run |
| 4 | 3 | 3,600 | A, A, A | "no progress" |
Without the repeat check this run would have used all 10 steps and 12,000 tokens to learn nothing. In the second run the queries differ, so the token budget (3,000) fires first, after 3,600 tokens.
Customer-support agents at scale always carry a per-conversation cost cap like this; a single looping conversation otherwise costs more than a thousand normal ones.
Follow-up questions to expect
- "A tool hangs for five minutes — which check catches it?" — None of them; they run between steps. Each tool call needs its own timeout, and the overall deadline should be passed into the tool.
- "How do you catch A, B, A, B loops?" — Count how often each fingerprint appears in the last N steps, not only consecutive repeats, and stop when any appears more than a threshold.
- "What do you return when a budget fires?" — A partial result with the reason, and ideally one last cheap model call asking it to summarise what it found so far, so the user gets something useful.