Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your planning agent performs well on short tasks but fails on workflows spanning hours or days. How do you build long-horizon AI agents that maintain goals, memory, and reliability over time?
What you need to know
Why long tasks break
Over hours, the conversation grows past the context window, gets summarised, and loses details. A deploy or crash kills the process and all its progress. Waiting two days for a manager's approval is impossible in a single running process. And the world changes while the agent waits.
The design
- The plan is data — a task tree in a database: each node has a status, dependencies, acceptance criteria and output artefacts.
- One step per turn — the agent reads the next ready node, does one step, writes the result back.
- Checkpoint every step — on a durable engine, so crashes, deploys and long waits are safe.
- Layered memory — working context (the current node), episodic memory (what was tried and what failed), semantic memory (learned facts with source and date).
- Re-ground on resume — re-read the goal and fetch current state from source systems; do not trust beliefs from yesterday.
- Budgets and check-ins — caps on steps, tokens and time per node; every N steps a supervisor asks "still moving toward the goal?" and escalates if not.
1from langgraph.checkpoint.postgres import PostgresSaver2from langgraph.types import interrupt34def request_approval(state):5 decision = interrupt({"approve": state["contract_draft"]}) # pauses; can wait for days6 return {"approved": decision == "yes"}78with PostgresSaver.from_conn_string(DB_URI) as checkpointer:9 checkpointer.setup()10 graph = builder.compile(checkpointer=checkpointer)11 config = {"configurable": {"thread_id": "vendor-onboarding-812"}}12 graph.invoke({"goal": "Onboard vendor 812"}, config)13 # days later, from any process, after the manager clicks approve:14 # graph.invoke(Command(resume="yes"), config)The checkpointer saves state to Postgres after each step, keyed by thread_id. interrupt pauses the run until a human answers; resuming uses Command from langgraph.types with the same thread_id, even from a different server.
Episodic memory is what stops the agent repeating a failed approach: "Tried vendor API on day 2; returned 403; ticket raised with IT."
A real-life example
Scenario, numbers made up. A procurement agent onboards new vendors: collect documents, verify GST and bank details, draft a contract, get legal approval, create the vendor in the ERP. It takes 3–5 days. In the first version, 60% of runs never finish: twice-weekly deploys kill in-flight runs, and resumed runs lose track of which documents were already received.
The team stores the plan as a task tree, runs the agent on LangGraph with a Postgres checkpointer, and uses interrupts for legal approval. On resume, the agent re-fetches the vendor's document status instead of trusting its notes. Completion rises to 91%, and the remaining failures escalate to a person with a clear record of what was tried.
Follow-up questions to expect
- "Temporal or LangGraph?" — Temporal is a general, very mature workflow engine; LangGraph is built around agent state and human-in-the-loop. Many teams run LangGraph agents inside larger workflows.
- "How do you evaluate long-horizon agents?" — Completion rate by task length, steps and cost per completed task, and human interventions per task.
- "What if the goal changes midway?" — Treat it as an update to the plan tree: re-plan the affected nodes and keep completed work that still applies.