Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your planning agent performs well on short tasks but fails on workflows spanning hours or days. How do you build long-horizon AI agents that maintain goals, memory, and reliability over time?


The agent as a loop over saved statere-read goaland world statepick nextready plan nodedo one stepwriteresult, checkpointwait:approval or retryon every resumesurvivesdeploysThe plan tree and the episodic log live in a database, not in the context window.
Long tasks stop failing when no single process has to stay alive, or remember, for the whole task.

What you need to know

Why long tasks break

Over hours, the conversation grows past the context window, gets summarised, and loses details. A deploy or crash kills the process and all its progress. Waiting two days for a manager's approval is impossible in a single running process. And the world changes while the agent waits.

The design

  1. The plan is data — a task tree in a database: each node has a status, dependencies, acceptance criteria and output artefacts.
  2. One step per turn — the agent reads the next ready node, does one step, writes the result back.
  3. Checkpoint every step — on a durable engine, so crashes, deploys and long waits are safe.
  4. Layered memory — working context (the current node), episodic memory (what was tried and what failed), semantic memory (learned facts with source and date).
  5. Re-ground on resume — re-read the goal and fetch current state from source systems; do not trust beliefs from yesterday.
  6. Budgets and check-ins — caps on steps, tokens and time per node; every N steps a supervisor asks "still moving toward the goal?" and escalates if not.
Python
from langgraph.checkpoint.postgres import PostgresSaverfrom langgraph.types import interruptdef request_approval(state):    decision = interrupt({"approve": state["contract_draft"]})   # pauses; can wait for days    return {"approved": decision == "yes"}with PostgresSaver.from_conn_string(DB_URI) as checkpointer:    checkpointer.setup()    graph = builder.compile(checkpointer=checkpointer)    config = {"configurable": {"thread_id": "vendor-onboarding-812"}}    graph.invoke({"goal": "Onboard vendor 812"}, config)    # days later, from any process, after the manager clicks approve:    # graph.invoke(Command(resume="yes"), config)

The checkpointer saves state to Postgres after each step, keyed by thread_id. interrupt pauses the run until a human answers; resuming uses Command from langgraph.types with the same thread_id, even from a different server.

Episodic memory is what stops the agent repeating a failed approach: "Tried vendor API on day 2; returned 403; ticket raised with IT."

A real-life example

Scenario, numbers made up. A procurement agent onboards new vendors: collect documents, verify GST and bank details, draft a contract, get legal approval, create the vendor in the ERP. It takes 3–5 days. In the first version, 60% of runs never finish: twice-weekly deploys kill in-flight runs, and resumed runs lose track of which documents were already received.

The team stores the plan as a task tree, runs the agent on LangGraph with a Postgres checkpointer, and uses interrupts for legal approval. On resume, the agent re-fetches the vendor's document status instead of trusting its notes. Completion rises to 91%, and the remaining failures escalate to a person with a clear record of what was tried.

Follow-up questions to expect

  • "Temporal or LangGraph?" — Temporal is a general, very mature workflow engine; LangGraph is built around agent state and human-in-the-loop. Many teams run LangGraph agents inside larger workflows.
  • "How do you evaluate long-horizon agents?" — Completion rate by task length, steps and cost per completed task, and human interventions per task.
  • "What if the goal changes midway?" — Treat it as an update to the plan tree: re-plan the affected nodes and keep completed work that still applies.