Agents & Tools Interview Prep

Course Content

Agents & Tools Interview Prep

6 sections · 40 lessons

How do you design state management for complex agent pipelines?


A restart in the middle of a multi-city bookingFlight booked,checkpoint savedServer restartsbefore the hotelNew workerloads RunStateFlight retryhits sameidempotency keyContinueswith the hotelWithout the key, the first version charged 61,000 rupees twice.
Resume means some step will run twice, so idempotency keys are what make a checkpoint safe rather than dangerous.

What you need to know

Three kinds of state

KindContentsProperties
ConversationMessages, tool calls, tool resultsAppend-only, grows, gets compacted
TaskGoal, plan, cursor, step status, budget, artefact refsSmall, structured, the source of truth for resume
ExternalBookings, payments, tickets, filesLives in real systems; the agent only holds IDs

A task-state object

Python
from typing import TypedDict, Literalclass Step(TypedDict):    id: int    tool: str    args: dict    status: Literal["pending", "running", "done", "failed", "awaiting_approval"]    result_ref: str | None          # key in blob storage, not the data itself    idempotency_key: strclass RunState(TypedDict):    run_id: str    user_id: str    goal: str    plan: list[Step]    cursor: int    budget_usd_left: float    messages_ref: str               # where the conversation is stored

Checkpoint after every step

Save RunState to a durable store (Postgres, Redis with persistence) after each step. This one habit gives you:

  • Crash recovery — a new worker loads the checkpoint and continues.
  • Long pauses — waiting hours for a human approval costs nothing.
  • Debugging — inspect the exact state before the bad step.
  • Replay — re-run from a checkpoint with a new prompt or model to test a fix.

LangGraph's checkpointer does this for graph state; with a hand-written loop, a table keyed by run_id is enough.

Idempotency: the rule that prevents double actions

If a worker crashes after create_transfer succeeds but before the checkpoint is written, the resumed run will call it again. Pass an idempotency key (for example run_id + step_id) to every side-effect API, so the second call returns the first result instead of acting again. Payment APIs support this widely; for your own tools, store the key with the result.

Keep payloads out

Store a 5 MB CSV or a 200-row query result in blob storage and keep only a reference in state. State stays small and fast to save, and the model gets a summary, not the raw data.

A real-life example

A travel-booking agent books multi-city trips: flights, then hotels, then a cab, then payment. In the first version, state was just the message list in memory.

An incident: the server restarted during a Mumbai–Dubai–London booking, after the flight was ticketed but before the hotel. The user retried, the agent started over, and the flight was booked twice — ₹61,000 charged twice before support reversed it.

The redesign:

  • RunState with a step list, saved after every step.
  • Each booking call sends idempotency_key = run_id:step_id; the airline API returns the existing ticket on a repeat.
  • Payment steps are awaiting_approval until the user confirms in the app; the run sleeps in the database, not in a process.
  • Search results are saved to storage; state keeps result_ref and a 5-line summary.

On the next restart, a new worker loaded the checkpoint, saw "flight: done, hotel: pending", and continued with the hotel.

Follow-up questions to expect

  • "Why not just store the messages and replay them?" — Replaying messages re-runs model calls, which may choose differently. Task state records what actually happened.
  • "Where does the conversation live if it gets compacted?" — Keep the full log in storage for audit, and a compacted version for the model.
  • "How do you handle a step that was 'running' at crash time?" — Treat it as unknown: check the external system (by idempotency key) to see if it happened, then mark it done or retry it.