LangGraph Agents

Course Content

LangGraph Agents

7 sections · 49 lessons

What are common failure modes, and how does persistence help?


What you need to know

FailureHow persistence helpsWhat else you need
Pod crash or deployResume with invoke(None, config)A job that finds unfinished threads
Provider 5xx / 429Retry one node; siblings' writes keptRetryPolicy with the right retry_on
Human takes daysPaused thread costs nothingExpiry policy
Runaway loopHistory shows which cycle repeatedCounters and recursion_limit
Wrong answerReplay or fork from before the mistakeEvals to catch it
Corrupted stateupdate_state to repair and continueAccess control on who can edit

What persistence does not fix

  • Committed side effects. A resumed node re-runs; without idempotency keys, emails and payments repeat.
  • Wrong logic. A bad route is saved and replayed exactly.
  • Bloat. Every extra key is written at every step.
  • Schema changes. Old checkpoints must still load with new code; renaming keys breaks paused threads.
  • Privacy. Checkpoints hold full conversations and tool results. Encrypt, restrict access, set retention, and honour deletion requests by deleting threads.

A real-life example

A health insurer reviewed six months of incidents on its claims graph:

  • 23 pod restarts mid-claim — all resumed from checkpoints, no customer impact.
  • A hospital-network API outage of 3 hours — claims paused at the verification node with a retry policy, then a fallback to manual verification.
  • One duplicate payout of Rs 42,000 — the payout node called the bank API, then crashed before its checkpoint; on resume it paid again. The fix was an idempotency key from claim_id, which the bank API checks.
  • A schema change renamed hospital_id to provider_id; 310 paused claims failed to resume until a migration node mapped the old key.

Persistence solved the first two; the last two needed engineering around it.

Follow-up questions to expect

  • "How do you make a payment node safe to re-run?" — Pass an idempotency key derived from stable data (claim id, refund id) so the provider ignores duplicates, or check "already paid?" before paying.
  • "How do you handle GDPR or DPDP deletion requests?" — Delete the user's threads with delete_thread and their Store namespaces, and make sure traces follow the same retention.
  • "What would you monitor?" — Threads stuck with a non-empty next and no interrupt, interrupt age, resume failures, and checkpoint size per thread.