Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Scenario – 5: Node Timeout Resilience


What you need to know

The scenario: one node calls a slow external API. Sometimes it hangs, and the whole graph run hangs with it.

Why one slow call hangs everything

Many HTTP clients wait for a very long time, or forever, by default. A node waiting on a dead connection blocks its superstep, and the graph cannot move on. Retries without a timeout make it worse: each retry waits forever too.

The four layers

  1. Timeout every call — e.g. httpx.Timeout(8.0) or asyncio.wait_for.
  2. Retry transient errors only — timeouts and connection errors, with backoff and a small cap.
  3. Catch and degrade — the node returns {"enrich_error": "timeout"}, and a conditional edge goes to a path that answers without enrichment.
  4. Checkpoint durably — a Postgres checkpointer saves state after each step, so a crashed run resumes where it stopped.
Python
import httpxfrom langgraph.types import RetryPolicyclient = httpx.AsyncClient(timeout=httpx.Timeout(8.0))async def enrich(state):    if state["deadline"] - now() < 3:                       # not enough budget left        return {"enrich_error": "skipped: time budget"}    try:        r = await client.get(f"{CRM}/customers/{state['customer_id']}")        r.raise_for_status()        return {"customer": r.json()}    except httpx.HTTPError as e:        return {"enrich_error": type(e).__name__}          # degrade, don't raisebuilder.add_node("enrich", enrich,    retry_policy=RetryPolicy(max_attempts=3, retry_on=(httpx.ConnectError,)))builder.add_conditional_edges("enrich",    lambda s: "answer_basic" if s.get("enrich_error") else "answer_full")

The retry policy covers connection errors raised before the node can catch them; everything else is turned into state. (Older LangGraph versions named this argument retry.)

Choosing what may degrade

DependencyRequired or optionalOn failure
Order database for "where is my order?"RequiredClear error with a retry option
CRM profile for personalisationOptionalAnswer without personalisation
Web search for extra contextOptionalAnswer from internal docs, say so

InMemorySaver (older name MemorySaver) keeps checkpoints in process memory, so it is for tests only; a crashed pod loses everything.

Metrics

Per-node p95 latency, per-node failure and retry rates, and the degraded-answer rate as a product metric. If 15% of answers are degraded, that is a conversation with the dependency's owner, not a graph fix.

A real-life example

Scenario, numbers made up. An insurance claims assistant enriches each chat with the customer's policy from a legacy system. On Monday mornings the legacy system slows down, and some runs hang for over two minutes until the load balancer kills them.

The team sets an 8-second timeout, retries connection errors up to three times, and adds a degraded path that asks the customer for their policy number instead. They move from MemorySaver to a Postgres checkpointer. p99 run time falls from over 120 seconds to 14 seconds, about 6% of Monday chats take the degraded path, and runs interrupted by deploys now resume instead of starting again.

Follow-up questions to expect

  • "Why not just retry more?" — Retries without timeouts multiply the hang, and retrying a dependency that is overloaded adds load to it. Cap and back off.
  • "How does the checkpointer help with timeouts?" — If the process dies, the run resumes from the last completed step, so earlier LLM calls are not paid for again.
  • "What is the time budget for?" — To keep total latency bounded when several things are slow at once: optional nodes skip themselves when little time is left.