Live Coding Interview Prep

Course Content

Live Coding Interview Prep

7 sections · 50 lessons

Write an agent loop that decides when to stop execution.


Fingerprints of a stuck agent's actionsAAA—01231,200 tokensthird repeat:no progressstep 4 never runsWith only a step limit, the same run would burn 10 steps and 12,000 tokens.
Checking the last three fingerprints stops a looping agent at step 3 instead of at the budget.

What you need to know

"Stop when the model says it is finished" is only the happy path. Agents fail by not stopping: they loop on the same search, wander, or burn the budget on a task they cannot do. A robust loop has two kinds of stop:

  • Success — the model emits a final answer.
  • Budgets and safety valves — steps, seconds, tokens, money, and "no progress". Each is a separate check because each catches a different failure.

Fingerprinting an action means turning it into a short, comparable string. json.dumps(action, sort_keys=True) gives the same text for the same action no matter the key order; hashing it keeps the stored value small.

time.monotonic() is a clock that only moves forward. time.time() follows the system clock, which NTP can adjust mid-run.

Python
import hashlib, json, timefrom collections import dequefrom collections.abc import Callablefrom dataclasses import dataclass, field@dataclassclass Budget:    max_steps: int = 10    max_seconds: float = 60.0    max_tokens: int = 50_000    max_usd: float = 0.50    repeat_limit: int = 3@dataclassclass StopController:    """Tracks spend and recent actions; should_stop() returns a reason or None."""    budget: Budget = field(default_factory=Budget)    steps: int = 0    tokens: int = 0    usd: float = 0.0    started: float = field(default_factory=time.monotonic)    recent: deque = field(default_factory=lambda: deque(maxlen=10))    def record(self, action: dict, tokens: int = 0, usd: float = 0.0) -> None:        self.steps, self.tokens, self.usd = self.steps + 1, self.tokens + tokens, self.usd + usd        key = json.dumps(action, sort_keys=True, default=str)        self.recent.append(hashlib.sha256(key.encode()).hexdigest()[:16])    def should_stop(self) -> str | None:        b = self.budget        if self.steps >= b.max_steps:            return "step limit"        if time.monotonic() - self.started > b.max_seconds:            return "time limit"        if self.tokens >= b.max_tokens:            return "token budget"        if self.usd >= b.max_usd:            return "cost budget"        last = list(self.recent)[-b.repeat_limit:]        if len(last) == b.repeat_limit and len(set(last)) == 1:            return "no progress: same action repeated"        return Nonedef agent_loop(question: str, policy: Callable, execute: Callable,               controller: StopController) -> dict:    """policy(transcript) -> action dict; execute(action) -> observation dict."""    transcript: list = [("question", question)]    while (reason := controller.should_stop()) is None:        action = policy(transcript)        if action.get("tool") == "final_answer":            return {"answer": action.get("answer"), "stopped": "completed", "steps": controller.steps}        observation = execute(action)        controller.record(action, tokens=observation.get("tokens", 0), usd=observation.get("usd", 0.0))        transcript.append((action, observation))    return {"answer": None, "stopped": reason, "steps": controller.steps}

The tricky parts:

  • The check runs before each step, so a budget that is already spent never triggers one more expensive call.
  • The repeat check looks at the last repeat_limit fingerprints only. Calling search("refund") three times in a row is stuck; calling it at steps 1 and 7 may be legitimate.
  • sort_keys=True makes {"a": 1, "b": 2} and {"b": 2, "a": 1} the same fingerprint.
  • The walrus := reads the reason and tests it in one line, so the final return can report which limit fired.

Complexity: record hashes one action, O(size of the action). should_stop is O(repeat_limit), which is constant. The deque(maxlen=10) keeps memory O(1) however long the run is.

A real-life example

A stuck policy that keeps searching for the same thing:

Python
stuck = lambda transcript: {"tool": "search", "args": {"query": "refund policy"}}execute = lambda action: {"result": "no results", "tokens": 1200}print(agent_loop("What is the refund policy?", stuck, execute, StopController()))# {'answer': None, 'stopped': 'no progress: same action repeated', 'steps': 3}tight = StopController(budget=Budget(max_tokens=3000))varied = lambda t: {"tool": "search", "args": {"query": f"refund try {len(t)}"}}print(agent_loop("What is the refund policy?", varied, execute, tight))# {'answer': None, 'stopped': 'token budget', 'steps': 3}

First run, step by step:

before stepstepstokenslast 3 fingerprintsshould_stop
100–None → run
211,200ANone → run
322,400A, ANone → run
433,600A, A, A"no progress"

Without the repeat check this run would have used all 10 steps and 12,000 tokens to learn nothing. In the second run the queries differ, so the token budget (3,000) fires first, after 3,600 tokens.

Customer-support agents at scale always carry a per-conversation cost cap like this; a single looping conversation otherwise costs more than a thousand normal ones.

Follow-up questions to expect

  • "A tool hangs for five minutes — which check catches it?" — None of them; they run between steps. Each tool call needs its own timeout, and the overall deadline should be passed into the tool.
  • "How do you catch A, B, A, B loops?" — Count how often each fingerprint appears in the last N steps, not only consecutive repeats, and stop when any appears more than a threshold.
  • "What do you return when a budget fires?" — A partial result with the reason, and ideally one last cheap model call asking it to summarise what it found so far, so the user gets something useful.