Course Content
AI Agent Frameworks
4 sections · 15 lessons
BabyAGI and Task Loop Management
Run a stock BabyAGI loop for twenty iterations on the objective "Write a five-part blog series on sustainable packaging" and watch the task queue length after each cycle. A real run looks like this:
iter 1 completed 1 task, created 3 queue: 3iter 5 completed 5 tasks, created 15 queue: 11iter 10 completed 10 tasks, created 29 queue: 20iter 15 completed 15 tasks, created 44 queue: 30iter 20 completed 20 tasks, created 58 queue: 39Twenty tasks done. Thirty-nine still waiting. The agent is further from finishing than when it started, and it has spent roughly 3 dollars getting there.
The arithmetic is unforgiving. Each iteration removes one task and adds k new ones, so after n iterations the queue holds
With k=3 and n=20: Q=1+20×2=41 — close to the 39 observed, the small difference being duplicate tasks that got filtered. The queue shrinks only when the average k is below 1. A loop that creates as many tasks as it completes never terminates, and a loop that creates more diverges. No amount of prompt improvement fixes that; it is a property of the recurrence.
A task-generating loop terminates if and only if the average number of new tasks per completed task falls below one. Everything else in the design is detail.
Queue, not stack — and why that changes everything
A goal-stack agent pops the most recently added item. That is depth-first: it takes one thread all the way down before looking at siblings. A task-loop agent like BabyAGI pops from a priority queue, which by default behaves breadth-first: it works across the whole problem at a shallow level before going deeper anywhere.
| Stack (depth-first) | Priority queue (breadth-first by default) | |
|---|---|---|
| Next item | Most recently added | Highest priority, oldest first among ties |
| Progress feels like | One branch completed, others untouched | Every branch advanced a little |
| Partial output if you stop early | One finished section | An outline of everything, nothing finished |
| Context reuse | Strong — siblings share a parent's findings | Weak — consecutive tasks may be unrelated |
| Suits | Deep research into one question | Broad coverage, reprioritisation as you learn |
The reason to accept the weaker context reuse is reprioritisation. A stack cannot reorder itself; whatever was pushed last runs next, even if what you just learned makes it irrelevant. A priority queue can be re-scored after every completed task, so a discovery in task 3 can promote task 17 to the front. That is the pattern's real selling point, and it is worthless if you never actually reprioritise — which most implementations do not.
The task data model
The single most common weakness in these loops is a task represented as a bare string. Give it structure:
1from dataclasses import dataclass, field2from enum import Enum3import itertools, time45_counter = itertools.count()67class TaskState(str, Enum):8 PENDING = "pending"; DONE = "done"; FAILED = "failed"; SKIPPED = "skipped"910@dataclass(order=True)11class Task:12 priority: int # 1 = do first, 9 = do last13 seq: int = field(default_factory=lambda: next(_counter), compare=True)14 description: str = field(default="", compare=False)15 created_by: str = field(default="root", compare=False)16 state: TaskState = field(default=TaskState.PENDING, compare=False)17 result: str = field(default="", compare=False)18 depth: int = field(default=0, compare=False)order=True with compare=True on only priority and seq means Python sorts tasks by priority first and insertion order second. That second key is not cosmetic: without it, heapq tries to compare two Task objects with equal priority, reaches description, and orders your work alphabetically. Tasks starting with "Analyse" run before tasks starting with "Write" for no reason anyone intended.
depth exists so that a task created by a task created by a task can be recognised and stopped. Without it there is no way to distinguish a fresh top-level task from a third-generation descendant.
The loop
1import heapq, json23class TaskLoopAgent:4 def __init__(self, llm, objective, max_iterations=15, max_queue=20,5 max_depth=2, token_budget=300_000):6 self.llm, self.objective = llm, objective7 self.max_iterations, self.max_queue = max_iterations, max_queue8 self.max_depth, self.token_budget = max_depth, token_budget9 self.queue: list[Task] = []10 self.done: list[Task] = []11 self.seen: set[str] = set()12 self.tokens = 01314 # ---------- 1. execute ----------15 def _execute(self, t: Task) -> str:16 context = "\n".join(f"- {d.description}: {d.result[:250]}"17 for d in self.done[-4:])18 return self._ask(19 f"Objective: {self.objective}\n"20 f"Recently completed:\n{context or '(nothing yet)'}\n\n"21 f"Task: {t.description}\n\n"22 "Complete this task concretely. Produce the actual artefact "23 "(text, list, numbers) - do not describe what you would do."24 )2526 # ---------- 2. create ----------27 def _create(self, t: Task) -> list[Task]:28 if t.depth >= self.max_depth:29 return []30 remaining = self.max_queue - len(self.queue)31 if remaining <= 0:32 return []33 allowance = min(2, remaining)3435 raw = self._ask(36 f"Objective: {self.objective}\n"37 f"Just completed: {t.description}\n"38 f"Result: {t.result[:900]}\n"39 f"Still queued: {[q.description for q in self.queue][:8]}\n\n"40 f"Create AT MOST {allowance} new tasks - fewer is better, and zero "41 "is correct if the objective is now adequately covered.\n"42 "Do NOT duplicate anything already queued or completed.\n"43 'Return JSON: [{"description": "...", "priority": 1-9}]'44 )45 try:46 items = json.loads(raw[raw.index("["): raw.rindex("]") + 1])47 except Exception:48 return []4950 fresh = []51 for it in items[:allowance]:52 key = it["description"].strip().lower()53 if key in self.seen:54 continue55 self.seen.add(key)56 fresh.append(Task(priority=int(it.get("priority", 5)),57 description=it["description"],58 created_by=t.description, depth=t.depth + 1))59 return fresh6061 # ---------- 3. reprioritise ----------62 def _reprioritise(self, learned: str) -> None:63 if not self.queue:64 return65 listing = "\n".join(f"{i}. [p{q.priority}] {q.description}"66 for i, q in enumerate(self.queue))67 raw = self._ask(68 f"Objective: {self.objective}\nNew information: {learned[:600]}\n\n"69 f"Queued tasks:\n{listing}\n\n"70 "Re-score priorities. 1 = do next, 9 = do last. A task made "71 "redundant by the new information should be 9.\n"72 'Return JSON: [{"index": 0, "priority": 3}, ...]'73 )74 try:75 for upd in json.loads(raw[raw.index("["): raw.rindex("]") + 1]):76 i = int(upd["index"])77 if 0 <= i < len(self.queue):78 self.queue[i].priority = max(1, min(9, int(upd["priority"])))79 except Exception:80 return81 heapq.heapify(self.queue)8283 # ---------- the cycle ----------84 def run(self, first_task: str) -> dict:85 heapq.heappush(self.queue, Task(priority=1, description=first_task))86 self.seen.add(first_task.strip().lower())8788 for i in range(self.max_iterations):89 if not self.queue or self.tokens >= self.token_budget:90 break91 t = heapq.heappop(self.queue)92 t.result = self._execute(t)93 t.state = TaskState.DONE94 self.done.append(t)9596 for nt in self._create(t):97 heapq.heappush(self.queue, nt)98 self._reprioritise(t.result)99100 print(f"iter {i+1:2d} done={len(self.done):2d} "101 f"queue={len(self.queue):2d} tokens={self.tokens:,}")102103 for q in self.queue:104 q.state = TaskState.SKIPPED105 return {"objective": self.objective, "completed": self.done,106 "unfinished": list(self.queue)}Where the convergence actually comes from
Four separate mechanisms push the average k below 1, and you need all four because each fails differently:
| Mechanism | Effect on k | Fails when |
|---|---|---|
allowance = min(2, remaining) | Caps k at 2 | Model reliably returns 2 — still divergent |
| "zero is correct if adequately covered" | Lets k reach 0 | Model is agreeable and rarely says zero |
max_depth | Forces k=0 at depth 2 | Broad shallow explosion before depth 2 |
max_queue | Hard stop, k=0 when full | Never — this is the backstop |
Work the numbers with these in place. Start with 1 task at depth 0. It generates 2 at depth 1. Each of those generates 2 at depth 2. Each depth-2 task generates none. Total tasks =1+2+4=7, all completed within 7 iterations, queue empty. Compare with the unguarded run: 20 iterations, 39 outstanding. The guards are the difference between a loop that finishes and one that only stops when you kill it.
Priority strategies
Asking the model for a priority number is the default and the weakest option — models cluster everything at 3 or 5, which makes the queue degenerate into first-in-first-out. Better strategies compute the number.
| Strategy | Rule | Good for | Watch out for |
|---|---|---|---|
| Model-assigned | Ask for 1–9 | Quick prototypes | Clustering; almost no discrimination |
| Dependency-first | Tasks nothing depends on get 9; blockers get 1 | Pipelines with real ordering | Needs a dependency graph |
| Depth-weighted | priority = 3 + depth | Keeping the loop shallow | Ignores importance entirely |
| Value-over-cost | ceil(9 - 8 * value / cost_estimate) | Fixed budgets | Cost estimates are guesses |
| Recency-decay | Add 1 to priority every 5 iterations queued | Preventing starvation | Can promote genuinely worthless tasks |
A hybrid that works well in practice: take the model's number, then adjust it mechanically.
1def effective_priority(t: Task, iteration: int) -> int:2 p = t.priority3 p += t.depth # deeper = later4 if any(w in t.description.lower()5 for w in ("research", "explore", "consider", "investigate")):6 p += 2 # vague work waits7 if any(w in t.description.lower()8 for w in ("draft", "write", "compute", "list", "publish")):9 p -= 1 # artefact-producing work first10 age = iteration - t.seq11 if age > 8:12 p -= 1 # anti-starvation13 return max(1, min(9, p))The verb heuristic is worth more than it looks. In a run scored this way, a task called "Research packaging regulations" gets priority 5+1+2=8 while "Draft part 1 on material choices" gets 5+1−1=5. The loop produces artefacts early and defers open-ended reading, which is exactly the behaviour you want if the run is cut short.
Memory: the deduplication problem
The seen set above catches exact string repeats and nothing else. Real duplicates do not repeat exactly:
iter 3 "Research eco-friendly packaging materials"iter 9 "Investigate sustainable packaging material options"iter 14 "Look into environmentally friendly packing materials"Three tasks, one job, three times the cost. Exact matching catches none of them. Embedding-based similarity catches all three:
1import numpy as np23class TaskMemory:4 def __init__(self, embed_fn, threshold=0.88):5 self.embed, self.threshold = embed_fn, threshold6 self.vectors, self.records = [], []78 def is_duplicate(self, text: str) -> str | None:9 v = self.embed(text)10 for vec, rec in zip(self.vectors, self.records):11 sim = float(np.dot(v, vec) /12 (np.linalg.norm(v) * np.linalg.norm(vec)))13 if sim >= self.threshold:14 return rec["result"]15 return None1617 def add(self, text: str, result: str) -> None:18 self.vectors.append(self.embed(text))19 self.records.append({"task": text, "result": result})Threshold choice is a real trade-off with real costs. At 0.95 you catch only near-identical phrasings and still pay for most duplicates. At 0.80 you start rejecting "Draft part 1" as a duplicate of "Draft part 2", which silently deletes work. Around 0.88 is a reasonable starting point, but the only honest way to set it is to log the pairs it rejects for a few runs and read them.
A run that finishes
Objective: "Write a five-part blog series on sustainable packaging." First task: "Draft the outline for all five parts." Limits: 15 iterations, queue 20, depth 2.
iter 1 [p1 d0] Draft the outline for all five parts created: "Draft part 1: material choices" (p2) "List 6 sources on packaging regulation" (p4) queue: 2iter 2 [p2 d1] Draft part 1: material choices created: "Draft part 2: supply chain impact" (p2) queue: 2iter 3 [p2 d2] Draft part 2: supply chain impact depth limit reached -> created 0 queue: 1iter 4 [p4 d1] List 6 sources on packaging regulation created: "Draft part 3: regulation" (p2) queue: 1iter 5 [p2 d2] Draft part 3: regulation created 0 queue: 0STOP: queue empty after 5 iterations. 5 tasks completed, 0 abandoned.tokens: 61,200 cost: 0.18 dollarsFive iterations rather than twenty, an empty queue rather than 39 outstanding, and a cost of 0.18 dollars rather than 3. But be honest about what was lost: parts 4 and 5 were never written. The depth limit cut off the branch that would have produced them.
That is the correct trade and it is worth stating plainly. A loop that stops cleanly with three of five parts drafted and an explicit list of what it did not do is a usable result. A loop that runs for an hour with all five parts half-started, twelve research tasks queued and no way to know where it stands is not. If you need all five parts, the fix is a first task that enumerates them explicitly — "Draft part 1" through "Draft part 5" as five sibling tasks at depth 1 — not a bigger depth limit.
"Done" in a task loop means the queue emptied, not that the objective was met. If you want the second meaning, you have to build the check that produces it.
Four ways the loop breaks
Everything gets the same priority
Ask a model for a 1–9 priority and you get 5, or 3, on almost everything. The queue becomes first-in-first-out and reprioritisation does nothing. Diagnose it by printing the priority histogram after a run; if more than half the tasks share one value, your priorities are decorative. Fix it by computing priority from features you control, as in effective_priority above.
Unbounded generation
The default failure, and the arithmetic is at the top of this lesson. "Create new tasks based on the result" with no number in the prompt produces three to five per iteration. Always state a maximum, always compute the allowance from remaining queue space, and always give the model explicit permission to return an empty list — models treat "create tasks" as an instruction to create tasks.
No iteration limit
An iteration cap is not the same as a queue cap. The queue can stay small while the loop churns forever on tasks that keep spawning single replacements. Cap iterations, tokens and wall-clock separately; each catches a different pathology.
Priority read backwards
Python's heapq is a min-heap: heappop returns the smallest value. So priority 1 runs first. Half of all implementations assume 9 means "most important", assign 9 to the critical path, and then wonder why the agent does the trivial work first while the important tasks sit at the back. The symptom is unmistakable once you know it: a run that produces perfect output in the last three iterations and noise in the first ten.
# If you prefer "9 = most important", negate on push:heapq.heappush(queue, (-importance, seq, task))Two beliefs to discard
"It will know when it is done." Nothing in the loop asks that question. Completion is a queue-empty condition, and the queue empties only when generation stops. If you want the agent to judge completeness, you must add an explicit check — a step that compares the objective against the completed set and answers "sufficient" or "not yet" — and give its answer the power to clear the queue. Without that, "done" means "out of budget".
"More tasks means more thorough." Beyond a point it means more redundant. In runs with unbounded generation, 20 to 30 per cent of tasks are near-duplicates of earlier ones and another chunk are vague reading tasks that produce no artefact. Thoroughness comes from the coverage of the first three or four tasks, not from the tail.
What to do with this pattern
Task loops are the right shape when the work is a set of loosely coupled items whose membership you discover as you go, and where reprioritising on new information genuinely matters — competitive sweeps, content pipelines, backlog triage. They are the wrong shape when you already know the items, because then you are paying a model to regenerate a list you could have written.
If you build one, instrument three numbers from the first run: tasks created per task completed (this is k, and it predicts termination), the priority histogram (flat means your priorities do nothing), and the proportion of completed tasks that produced a concrete artefact rather than a description of work to be done. That third number is the one that separates a loop that produced a blog series from a loop that produced a plan to produce a blog series — and it is the one nobody measures until they read a run's output and find nothing in it they can publish.