LangGraph Agents

Course Content

LangGraph Agents

7 sections · 49 lessons

What is checkpointing in LangGraph, and what gets saved at checkpoints?


The checkpoint chain for refund OD-88123inputafterlookup_orderafter scoreafter approveafterissue_refundnullnext: approve,interrupt pendingnext is empty: doneThe checkpoint after score still holds fraud_score 0.07 for the auditor.
The auditor's question two weeks later was answered by reading one saved checkpoint, not by re-running anything.

What you need to know

Python
from langgraph.checkpoint.memory import InMemorySavergraph = builder.compile(checkpointer=InMemorySaver())config = {"configurable": {"thread_id": "refund-OD-88123"}}graph.invoke({"order_id": "OD-88123", "amount_inr": 8500}, config)snap = graph.get_state(config)print(snap.values)        # the stateprint(snap.next)          # nodes scheduled next, () if finishedprint(snap.metadata)      # step number, sourceprint(snap.interrupts)    # pending interrupt payloadsprint(snap.parent_config) # the previous checkpoint

What is inside a checkpoint

PartWhy it exists
Channel values (the state)Continue from here without recomputing
nextKnow which nodes to run on resume
Pending writesIf one parallel node fails, the others' results are kept and not re-run
Metadata (step, source)Debugging and history listing
InterruptsWhat a paused run is waiting for
checkpoint_id and parent idA chain you can walk back through

The StateSnapshot returned by get_state exposes these as values, next, config, metadata, created_at, parent_config, tasks and interrupts.

Durability modes

Pass durability= to invoke or stream:

  • "sync" — write each checkpoint before the next step starts. Safest, slowest.
  • "async" (default) — write while the next step runs. Small risk of losing the last step in a crash.
  • "exit" — write only when the run ends, errors or interrupts. Fastest, but no mid-run recovery.

Serialisation

Checkpoints are serialised (JSON-like with a fallback for other Python types). Everything in state must be serialisable. For sensitive data, EncryptedSerializer can encrypt checkpoints at rest.

A real-life example

A refund graph has four steps: lookup_order, score, approve, issue_refund. For order OD-88123 (Rs 8,500) the history holds a checkpoint for the input, one after lookup_order, and one after score — that last one shows next=("approve",) and an interrupt payload asking a reviewer. When the reviewer approves two hours later, the run continues from that saved point and adds a checkpoint after approve and one after issue_refund. An auditor later asks "what fraud score did we see before approving?" — the answer is in the checkpoint after score: 0.07.

Follow-up questions to expect

  • "Is a checkpoint saved after every node?" — After every super-step. Parallel nodes in one step share one checkpoint.
  • "How big can a checkpoint get?" — As big as your state; the checkpointer stores changed channels efficiently, but large lists still grow. Keep blobs out of state.
  • "Does checkpointing slow the graph?" — Slightly. With durability="async" the write overlaps the next step; for short, low-value runs, "exit" or no checkpointer at all is fine.