Course Content
Agents & Tools Interview Prep
6 sections · 40 lessons
How do you design state management for complex agent pipelines?
What you need to know
Three kinds of state
| Kind | Contents | Properties |
|---|---|---|
| Conversation | Messages, tool calls, tool results | Append-only, grows, gets compacted |
| Task | Goal, plan, cursor, step status, budget, artefact refs | Small, structured, the source of truth for resume |
| External | Bookings, payments, tickets, files | Lives in real systems; the agent only holds IDs |
A task-state object
1from typing import TypedDict, Literal23class Step(TypedDict):4 id: int5 tool: str6 args: dict7 status: Literal["pending", "running", "done", "failed", "awaiting_approval"]8 result_ref: str | None # key in blob storage, not the data itself9 idempotency_key: str1011class RunState(TypedDict):12 run_id: str13 user_id: str14 goal: str15 plan: list[Step]16 cursor: int17 budget_usd_left: float18 messages_ref: str # where the conversation is storedCheckpoint after every step
Save RunState to a durable store (Postgres, Redis with persistence) after each step. This one habit gives you:
- Crash recovery — a new worker loads the checkpoint and continues.
- Long pauses — waiting hours for a human approval costs nothing.
- Debugging — inspect the exact state before the bad step.
- Replay — re-run from a checkpoint with a new prompt or model to test a fix.
LangGraph's checkpointer does this for graph state; with a hand-written loop, a table keyed by run_id is enough.
Idempotency: the rule that prevents double actions
If a worker crashes after create_transfer succeeds but before the checkpoint is written, the resumed run will call it again. Pass an idempotency key (for example run_id + step_id) to every side-effect API, so the second call returns the first result instead of acting again. Payment APIs support this widely; for your own tools, store the key with the result.
Keep payloads out
Store a 5 MB CSV or a 200-row query result in blob storage and keep only a reference in state. State stays small and fast to save, and the model gets a summary, not the raw data.
A real-life example
A travel-booking agent books multi-city trips: flights, then hotels, then a cab, then payment. In the first version, state was just the message list in memory.
An incident: the server restarted during a Mumbai–Dubai–London booking, after the flight was ticketed but before the hotel. The user retried, the agent started over, and the flight was booked twice — ₹61,000 charged twice before support reversed it.
The redesign:
RunStatewith a step list, saved after every step.- Each booking call sends
idempotency_key = run_id:step_id; the airline API returns the existing ticket on a repeat. - Payment steps are
awaiting_approvaluntil the user confirms in the app; the run sleeps in the database, not in a process. - Search results are saved to storage; state keeps
result_refand a 5-line summary.
On the next restart, a new worker loaded the checkpoint, saw "flight: done, hotel: pending", and continued with the hotel.
Follow-up questions to expect
- "Why not just store the messages and replay them?" — Replaying messages re-runs model calls, which may choose differently. Task state records what actually happened.
- "Where does the conversation live if it gets compacted?" — Keep the full log in storage for audit, and a compacted version for the model.
- "How do you handle a step that was 'running' at crash time?" — Treat it as unknown: check the external system (by idempotency key) to see if it happened, then mark it done or retry it.