Course Content
AI Agent Fundamentals
5 sections · 13 lessons
Agent Architectures (Reactive, Deliberative, Hybrid)
A warehouse robot is carrying a pallet down aisle 7. Its planner is excellent: it has a map of all 12,000 square metres, it knows where every rack is, and it computes a provably shortest route to the loading bay. The computation takes 2.8 seconds.
1.1 seconds into that computation, a human steps into aisle 7.
The robot does not stop. It cannot stop, because "stop" is not a thought it is currently having — it is 40% of the way through a graph search over a world model that no longer matches the world. When the plan finally comes out at 2.8 seconds, it is a beautiful route through a space that now contains a person.
Now swap it for a robot with no planner at all. Its entire program is a list of conditions: if obstacle within 50cm, stop. If path clear, advance. If at bay, release pallet. Reaction time: 30 milliseconds. It stops instantly.
It also takes 40 minutes to reach the loading bay, because it has no map and wanders. Point it at a task requiring three coordinated steps and it fails completely.
Neither robot is badly built. They are built on opposite architectures, and each one's strength is precisely the other's weakness. Understanding that trade — and how to get both — is what this lesson is about.
What "architecture" means here
An agent architecture is the answer to one question: how much thinking happens between a percept arriving and an action leaving?
At one extreme, none — the percept indexes directly into a response. At the other extreme, a great deal — the percept updates a world model, which feeds a planner, which produces a sequence of actions. Everything else sits between.
REACTIVE percept ──▶ rule match ──▶ action (fast, no memory, no lookahead)DELIBERATIVE percept ──▶ update world model ──▶ plan ──▶ action (slow, needs a model, produces sequences)HYBRID percept ──┬─▶ reflex layer ──────────▶ action └─▶ planner ──▶ intentions ─▶ action (reflexes preempt; planning runs alongside)Architecture is a choice about where you spend time. Reactive agents spend it at design time, writing rules. Deliberative agents spend it at run time, searching. Hybrids spend a little of both and pay in complexity.
Reactive agents
A reactive agent maps the current percept directly to an action. No internal model of the world, no memory of what happened before, no prediction of what happens next. Formally it is a function action = f(percept), and that is genuinely all of it.
The canonical example
A thermostat. Target 21°C, hysteresis band of 0.5°C to stop it flapping.
1def thermostat(temp_c, target=21.0, band=0.5):2 if temp_c < target - band: # below 20.53 return "HEAT_ON"4 if temp_c > target + band: # above 21.55 return "HEAT_OFF"6 return "HOLD"Trace it: at 19.8°C → HEAT_ON. At 20.9°C → HOLD (inside the band, and note it does not care that it was heating a moment ago). At 21.7°C → HEAT_OFF. The hysteresis band matters: without it, at exactly 21.0°C with sensor noise of ±0.2°C the device would toggle several times a minute and burn out its relay. That band is a design-time insight baked into a rule — which is the reactive pattern in miniature.
Reactive in an LLM system
Reactive does not mean primitive. A production intent router is reactive:
1RULES = [2 (r"\b(refund|money back|charge ?back)\b", "refund_flow"),3 (r"\b(where.{0,15}order|track|shipping)\b", "tracking_flow"),4 (r"\b(cancel|unsubscribe|close account)\b", "cancellation_flow"),5 (r"\b(password|log ?in|2fa|locked out)\b", "auth_flow"),6]78def route(message):9 text = message.lower()10 for pattern, handler in RULES:11 if re.search(pattern, text):12 return handler13 return "general_llm_flow"Latency: under a millisecond. Cost: zero. Behaviour: perfectly predictable, and testable with a table of inputs and expected outputs. For the 60–70% of support traffic that is one of four things, this beats an LLM on every axis that matters.
Where it breaks
Reactive agents fail on anything requiring history or lookahead. Watch the same router meet a real conversation:
User: I want to cancel my order. -> cancellation_flow ✓Agent: Which order?User: The blue one. -> general_llm_flow ✗ (no memory of "cancel")User: Actually where is it first? -> tracking_flow ✓Agent: In transit, arriving Friday.User: Fine, then cancel it. -> cancellation_flow ✓ (but which order? gone)Every individual routing decision is correct in isolation. The conversation is still broken, because a stateless function cannot hold a thread. Add memory to fix it and you have stopped being reactive.
| Reactive: strengths | Reactive: weaknesses |
|---|---|
| Millisecond latency | No memory — cannot follow a conversation |
| Trivial to test exhaustively | No lookahead — cannot sequence actions |
| Cannot loop or thrash | Rule count explodes as cases multiply |
| Fully auditable: you can read every rule | Silent on anything unanticipated |
| Cheap — often no model call at all | Cannot improve without a human editing rules |
Use reactive when the input space is small and known, the response is fixed, latency is critical, or the action is a safety reflex that must never wait on a planner.
Deliberative agents
A deliberative agent maintains an explicit model of the world, considers possible futures, and produces a plan — a sequence of actions expected to reach the goal — before acting.
percepts ──▶ [ WORLD MODEL ] what is true now │ ▼ [ GOAL STATE ] what should be true │ ▼ [ PLANNER ] search over action sequences │ ▼ [ PLAN: a1, a2, a3, ... ] │ ▼ execute a1, a2, a3 ...Worked example: booking a trip
Goal: be in Tokyo by 09:00 on 3 May, spend under 90,000 rupees, and do not travel overnight on a working day.
The agent's world model holds flight options, hotel availability, the user's calendar and the budget. The planner searches sequences:
Candidate A: fly 2 May 14:00 (52,000) + hotel 2 nights (24,000) = 76,000, arrives 2 May 23:00, sleep, meeting 09:00 ✓Candidate B: fly 3 May 01:30 (38,000) + hotel 1 night (12,000) = 50,000, but departs during Fri night ✗ violates constraint 3Candidate C: fly 1 May 09:00 (61,000) + hotel 3 nights (36,000) = 97,000 ✗ over budgetChosen: A. Total 76,000, 14,000 under budget, no overnight leg.Only a deliberative agent can do this. Candidate B is cheaper by 26,000 and a reactive "pick the cheapest flight" rule would take it — and violate a constraint the user cares about. Evaluating a whole sequence against all constraints before committing is exactly the capability that planning buys.
Where it breaks
The world changes while you plan. This is the warehouse robot, and it is the fundamental problem. If your plan takes 3 seconds to compute and the environment changes every 1 second, your plans are always describing a world that has passed.
The model is wrong. A plan is only as good as the world model behind it. If the agent believes flight XY tickets are available and they sold out an hour ago, every step after that assumption is fiction. Deliberative agents fail confidently, because internal consistency feels like correctness.
Search cost grows brutally. With a branching factor of 5 and a depth of 3 you examine 125 sequences — fine. Depth 8 is 390,625. Depth 12 is over 244 million. Add one more option per step and depth 12 becomes 2.2 billion. Long-horizon planning without a good heuristic to prune with is not slow; it is impossible.
| Deliberative: strengths | Deliberative: weaknesses |
|---|---|
| Handles multi-step, constrained goals | Latency measured in seconds |
| Can prove a plan satisfies constraints | Requires an accurate world model |
| Can compare alternatives before committing | Stale plans in dynamic environments |
| Explains itself — the plan is the explanation | Search cost explodes with horizon |
| Generalises to unseen goal combinations | Brittle when a mid-plan step fails |
Use deliberative when actions have ordering dependencies, when constraints must be checked across the whole sequence, and when the environment is slow enough that a plan survives long enough to execute.
Hybrid architectures
The warehouse robot's problem has an obvious fix once stated: let the reflex interrupt the planner. That is the hybrid idea. Run a fast reactive layer that owns anything urgent, and a slow deliberative layer that owns anything that needs thought, with a rule for who wins.
percept │ ├──────────────▶ REFLEX LAYER (µs–ms) │ obstacle? stop. fire? evacuate. │ │ │ └── if triggered: ACT NOW, preempt │ └──────────────▶ DELIBERATIVE LAYER (100ms–s) update model, plan, refine intentions │ └── if no reflex fired: ACT on planThe layering rule that makes this work: the reactive layer always has priority, and the deliberative layer must be interruptible. If a plan cannot be abandoned mid-execution, you have two architectures stapled together, not a hybrid.
BDI: the standard way to organise the thinking layer
BDI stands for Belief–Desire–Intention, and it is the most widely used framing for hybrid agents. Three stores:
- Beliefs — what the agent takes to be true about the world right now. Updated by perception. Beliefs can be wrong; that is the point of calling them beliefs rather than facts.
- Desires — goals the agent would like to achieve. There can be many, and they can conflict ("resolve fast" and "never break policy" pull against each other).
- Intentions — the subset of desires the agent has committed to and is actively pursuing. Commitment is the key idea: an intention persists across steps, so the agent does not abandon a goal the instant something else looks marginally more attractive.
Why commitment matters. Without it, an agent with desires "answer the customer" and "check for new tickets" ping-pongs — one step towards each, forever, finishing neither. Intentions give it the stubbornness to see something through. Too much stubbornness and it pursues a goal after the reason for it evaporated, so the deliberation cycle includes a reconsideration test.
A BDI support agent
1class BDIAgent:2 def __init__(self):3 self.beliefs = {} # facts about the world4 self.desires = [] # (goal, priority)5 self.intention = None # the one we are committed to67 def perceive(self, event):8 self.beliefs.update(event)910 def deliberate(self):11 # 1. Generate options consistent with current beliefs12 options = [(g, p) for g, p in self.desires if self.feasible(g)]1314 # 2. Should we reconsider? Only if the current intention is15 # finished, impossible, or clearly outranked.16 if self.intention:17 still_ok = (self.feasible(self.intention)18 and not self.achieved(self.intention))19 better = max(options, key=lambda o: o[1], default=(None, -1))20 outranked = better[1] > self.priority_of(self.intention) + 121 if still_ok and not outranked:22 return self.intention # stay committed2324 # 3. Commit to the highest-priority feasible option25 self.intention = max(options, key=lambda o: o[1])[0] if options \26 else None27 return self.intention2829 def step(self, event):30 self.perceive(event)31 if reflex := self.check_reflexes(): # reactive layer wins32 return reflex33 goal = self.deliberate()34 return self.next_action_for(goal)3536 def check_reflexes(self):37 if self.beliefs.get("customer_says_legal"):38 return "escalate_to_legal" # never planned around39 if self.beliefs.get("pii_detected"):40 return "redact_and_halt"41 return NoneRead the outranked line carefully — better[1] > priority + 1, not > priority. That margin is the anti-dithering device. A competing goal must be meaningfully better, not marginally better, to steal commitment. Set the margin to zero and the agent flip-flops on noise; set it too high and it ignores genuine emergencies. This one constant is where a lot of BDI tuning lives.
And note that check_reflexes runs before deliberate and returns directly. A legal threat does not enter the planner as a high-priority desire to be weighed. It bypasses thinking entirely, which is exactly what you want for anything that must never be reasoned away.
| Hybrid: strengths | Hybrid: weaknesses |
|---|---|
| Fast where speed matters, thoughtful elsewhere | Two systems to build, test and reason about |
| Safety reflexes cannot be planned around | Layer-interaction bugs are subtle and rare |
| Degrades gracefully — reflexes survive planner failure | Arbitration rules need real tuning |
| Commitment prevents goal thrashing | Harder to explain a decision — which layer chose it? |
Side by side
| Reactive | Deliberative | Hybrid / BDI | |
|---|---|---|---|
| World model | None | Explicit, central | Beliefs, possibly stale |
| Memory | None | Plan state | Beliefs + intentions |
| Lookahead | Zero | Multi-step | Multi-step, interruptible |
| Response time | Microseconds | 100 ms – seconds | Reflex µs, plan ~seconds |
| Handles novelty | Badly — falls through rules | Well, if the model covers it | Well |
| Dynamic environment | Excellent | Poor — plans go stale | Good |
| Debuggability | Trivial | Moderate — read the plan | Hard |
| Build effort | Hours | Days | Weeks |
| Compute per decision | Negligible | High | Moderate to high |
| Typical use | Routing, safety cut-offs, filters | Trip planning, research, scheduling | Robotics, trading, live support |
Choosing one
Four questions, answered in order. The first "yes" that forces an upgrade wins.
1. Does the correct action depend on anything other than the current input? NO ──▶ REACTIVE. Stop here. Do not build more. YES ──▶ continue2. Do actions have to happen in a particular order, or must constraints hold across a whole sequence? NO ──▶ REACTIVE + a memory store is usually enough YES ──▶ continue3. Does the environment change faster than you can plan, OR is there any action that must happen immediately and must never be reasoned away? NO ──▶ DELIBERATIVE YES ──▶ HYBRID4. (Sanity check) Can you afford to build, test and operate two coordinated layers? NO ──▶ DELIBERATIVE + a hard-coded safety interlock outside the agent. Simpler, and covers most of the benefit.| If your system... | Build |
|---|---|
| Classifies or routes one input at a time | Reactive |
| Answers questions from a fixed knowledge base | Reactive with retrieval |
| Must satisfy a budget and a deadline together | Deliberative |
| Researches a topic across many sources | Deliberative |
| Controls something physical | Hybrid, no exceptions |
| Trades, deploys, or moves money | Hybrid — the reflex layer is your circuit breaker |
| Holds a multi-turn conversation and takes actions | Hybrid, lightweight |
Choose the simplest architecture that can express the failure you are most afraid of. If your worst case is "wrong answer", deliberative is enough. If it is "acted too late", you need reflexes.
Where people get this wrong
Reaching for deliberation because the task sounds sophisticated. "Intelligent email triage" sounds like it needs planning. It is a classifier. Building a planner for it means paying seconds of latency and dollars of tokens to produce a one-word answer, with a new failure mode — the loop — that the classifier did not have.
Putting safety rules in the planner. A prompt saying "never issue a refund above 500 dollars" is a desire competing with other desires. Given enough pressure in a conversation, the model will construct a reason why this case is different. Safety rules belong in the reflex layer or, better, outside the agent in the service that executes the action, where no reasoning can reach them.
Building a hybrid where the layers never actually interact. A common shape: a reflex layer that checks three conditions which have never once fired in production, plus a planner doing all the work. That is a deliberative agent with extra code. Either the reflexes are load-bearing or delete them.
Letting the planner replan on every percept. This is BDI without intentions. The agent recomputes from scratch each step, picks a slightly different plan because the model is stochastic, and makes no progress on any of them. Commitment is not an optimisation; it is what makes multi-step behaviour possible.
What this means when you build one
Most real LLM systems end up as light hybrids, and they get there by accident and badly. You will build a deliberative agent because that is what the frameworks encourage, then discover it needs to never do a particular thing, and then bolt on checks. Do it deliberately instead.
Start by writing down the actions that must never be reasoned about — the escalations, the hard limits, the kill switches. Implement those as straight conditionals that run before the model call and return immediately. That is your reactive layer, and it should be small enough to read in one screen.
Then build the deliberative part for everything else, and make it interruptible: a step budget, a cancellation check between steps, and no operation so long that a reflex cannot preempt it.
If your agent holds a goal across multiple turns, add intentions — one field recording what it is currently committed to, and an explicit test for when to abandon it. Include the margin. An agent that switches goals only when something is clearly better looks far more competent than one that switches whenever something is slightly better, and the difference between those two behaviours is a single comparison operator.