Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Scenario – 5: Node Timeout Resilience
What you need to know
The scenario: one node calls a slow external API. Sometimes it hangs, and the whole graph run hangs with it.
Why one slow call hangs everything
Many HTTP clients wait for a very long time, or forever, by default. A node waiting on a dead connection blocks its superstep, and the graph cannot move on. Retries without a timeout make it worse: each retry waits forever too.
The four layers
- Timeout every call — e.g.
httpx.Timeout(8.0)orasyncio.wait_for. - Retry transient errors only — timeouts and connection errors, with backoff and a small cap.
- Catch and degrade — the node returns
{"enrich_error": "timeout"}, and a conditional edge goes to a path that answers without enrichment. - Checkpoint durably — a Postgres checkpointer saves state after each step, so a crashed run resumes where it stopped.
1import httpx2from langgraph.types import RetryPolicy34client = httpx.AsyncClient(timeout=httpx.Timeout(8.0))56async def enrich(state):7 if state["deadline"] - now() < 3: # not enough budget left8 return {"enrich_error": "skipped: time budget"}9 try:10 r = await client.get(f"{CRM}/customers/{state['customer_id']}")11 r.raise_for_status()12 return {"customer": r.json()}13 except httpx.HTTPError as e:14 return {"enrich_error": type(e).__name__} # degrade, don't raise1516builder.add_node("enrich", enrich,17 retry_policy=RetryPolicy(max_attempts=3, retry_on=(httpx.ConnectError,)))18builder.add_conditional_edges("enrich",19 lambda s: "answer_basic" if s.get("enrich_error") else "answer_full")The retry policy covers connection errors raised before the node can catch them; everything else is turned into state. (Older LangGraph versions named this argument retry.)
Choosing what may degrade
| Dependency | Required or optional | On failure |
|---|---|---|
| Order database for "where is my order?" | Required | Clear error with a retry option |
| CRM profile for personalisation | Optional | Answer without personalisation |
| Web search for extra context | Optional | Answer from internal docs, say so |
InMemorySaver (older name MemorySaver) keeps checkpoints in process memory, so it is for tests only; a crashed pod loses everything.
Metrics
Per-node p95 latency, per-node failure and retry rates, and the degraded-answer rate as a product metric. If 15% of answers are degraded, that is a conversation with the dependency's owner, not a graph fix.
A real-life example
Scenario, numbers made up. An insurance claims assistant enriches each chat with the customer's policy from a legacy system. On Monday mornings the legacy system slows down, and some runs hang for over two minutes until the load balancer kills them.
The team sets an 8-second timeout, retries connection errors up to three times, and adds a degraded path that asks the customer for their policy number instead. They move from MemorySaver to a Postgres checkpointer. p99 run time falls from over 120 seconds to 14 seconds, about 6% of Monday chats take the degraded path, and runs interrupted by deploys now resume instead of starting again.
Follow-up questions to expect
- "Why not just retry more?" — Retries without timeouts multiply the hang, and retrying a dependency that is overloaded adds load to it. Cap and back off.
- "How does the checkpointer help with timeouts?" — If the process dies, the run resumes from the last completed step, so earlier LLM calls are not paid for again.
- "What is the time budget for?" — To keep total latency bounded when several things are slow at once: optional nodes skip themselves when little time is left.