Course Content
LangGraph Agents
7 sections · 49 lessons
What are common failure modes, and how does persistence help?
What you need to know
| Failure | How persistence helps | What else you need |
|---|---|---|
| Pod crash or deploy | Resume with invoke(None, config) | A job that finds unfinished threads |
| Provider 5xx / 429 | Retry one node; siblings' writes kept | RetryPolicy with the right retry_on |
| Human takes days | Paused thread costs nothing | Expiry policy |
| Runaway loop | History shows which cycle repeated | Counters and recursion_limit |
| Wrong answer | Replay or fork from before the mistake | Evals to catch it |
| Corrupted state | update_state to repair and continue | Access control on who can edit |
What persistence does not fix
- Committed side effects. A resumed node re-runs; without idempotency keys, emails and payments repeat.
- Wrong logic. A bad route is saved and replayed exactly.
- Bloat. Every extra key is written at every step.
- Schema changes. Old checkpoints must still load with new code; renaming keys breaks paused threads.
- Privacy. Checkpoints hold full conversations and tool results. Encrypt, restrict access, set retention, and honour deletion requests by deleting threads.
A real-life example
A health insurer reviewed six months of incidents on its claims graph:
- 23 pod restarts mid-claim — all resumed from checkpoints, no customer impact.
- A hospital-network API outage of 3 hours — claims paused at the verification node with a retry policy, then a fallback to manual verification.
- One duplicate payout of Rs 42,000 — the payout node called the bank API, then crashed before its checkpoint; on resume it paid again. The fix was an idempotency key from
claim_id, which the bank API checks. - A schema change renamed
hospital_idtoprovider_id; 310 paused claims failed to resume until a migration node mapped the old key.
Persistence solved the first two; the last two needed engineering around it.
Follow-up questions to expect
- "How do you make a payment node safe to re-run?" — Pass an idempotency key derived from stable data (claim id, refund id) so the provider ignores duplicates, or check "already paid?" before paying.
- "How do you handle GDPR or DPDP deletion requests?" — Delete the user's threads with
delete_threadand their Store namespaces, and make sure traces follow the same retention. - "What would you monitor?" — Threads stuck with a non-empty
nextand no interrupt, interrupt age, resume failures, and checkpoint size per thread.