AI Agent Fundamentals

Agent Architectures (Reactive, Deliberative, Hybrid)


A warehouse robot is carrying a pallet down aisle 7. Its planner is excellent: it has a map of all 12,000 square metres, it knows where every rack is, and it computes a provably shortest route to the loading bay. The computation takes 2.8 seconds.

1.1 seconds into that computation, a human steps into aisle 7.

The robot does not stop. It cannot stop, because "stop" is not a thought it is currently having — it is 40% of the way through a graph search over a world model that no longer matches the world. When the plan finally comes out at 2.8 seconds, it is a beautiful route through a space that now contains a person.

Now swap it for a robot with no planner at all. Its entire program is a list of conditions: if obstacle within 50cm, stop. If path clear, advance. If at bay, release pallet. Reaction time: 30 milliseconds. It stops instantly.

It also takes 40 minutes to reach the loading bay, because it has no map and wanders. Point it at a task requiring three coordinated steps and it fails completely.

Neither robot is badly built. They are built on opposite architectures, and each one's strength is precisely the other's weakness. Understanding that trade — and how to get both — is what this lesson is about.

The hybrid stack in the warehouse robotReflex layer —stop, in millisecondsSequencing —resume or rerouteDeliberativeplanner — 2.8 secondstopbottomA plan that is provably shortest is still wrong if it arrives after the collision.
Hybrids exist because urgency and quality live on different clocks: reflexes below, deliberation above.

What "architecture" means here

An agent architecture is the answer to one question: how much thinking happens between a percept arriving and an action leaving?

At one extreme, none — the percept indexes directly into a response. At the other extreme, a great deal — the percept updates a world model, which feeds a planner, which produces a sequence of actions. Everything else sits between.

Text
REACTIVE          percept ──▶ rule match ──▶ action                  (fast, no memory, no lookahead)DELIBERATIVE      percept ──▶ update world model ──▶ plan ──▶ action                  (slow, needs a model, produces sequences)HYBRID            percept ──┬─▶ reflex layer ──────────▶ action                            └─▶ planner ──▶ intentions ─▶ action                  (reflexes preempt; planning runs alongside)

Architecture is a choice about where you spend time. Reactive agents spend it at design time, writing rules. Deliberative agents spend it at run time, searching. Hybrids spend a little of both and pay in complexity.

Reactive agents

A reactive agent maps the current percept directly to an action. No internal model of the world, no memory of what happened before, no prediction of what happens next. Formally it is a function action = f(percept), and that is genuinely all of it.

The canonical example

A thermostat. Target 21°C, hysteresis band of 0.5°C to stop it flapping.

Python
def thermostat(temp_c, target=21.0, band=0.5):    if temp_c < target - band:      # below 20.5        return "HEAT_ON"    if temp_c > target + band:      # above 21.5        return "HEAT_OFF"    return "HOLD"

Trace it: at 19.8°C → HEAT_ON. At 20.9°C → HOLD (inside the band, and note it does not care that it was heating a moment ago). At 21.7°C → HEAT_OFF. The hysteresis band matters: without it, at exactly 21.0°C with sensor noise of ±0.2°C the device would toggle several times a minute and burn out its relay. That band is a design-time insight baked into a rule — which is the reactive pattern in miniature.

Reactive in an LLM system

Reactive does not mean primitive. A production intent router is reactive:

Python
RULES = [    (r"\b(refund|money back|charge ?back)\b",   "refund_flow"),    (r"\b(where.{0,15}order|track|shipping)\b", "tracking_flow"),    (r"\b(cancel|unsubscribe|close account)\b", "cancellation_flow"),    (r"\b(password|log ?in|2fa|locked out)\b",  "auth_flow"),]def route(message):    text = message.lower()    for pattern, handler in RULES:        if re.search(pattern, text):            return handler    return "general_llm_flow"

Latency: under a millisecond. Cost: zero. Behaviour: perfectly predictable, and testable with a table of inputs and expected outputs. For the 60–70% of support traffic that is one of four things, this beats an LLM on every axis that matters.

Where it breaks

Reactive agents fail on anything requiring history or lookahead. Watch the same router meet a real conversation:

Text
User: I want to cancel my order.       -> cancellation_flow  ✓Agent: Which order?User: The blue one.                    -> general_llm_flow   ✗                                          (no memory of "cancel")User: Actually where is it first?      -> tracking_flow      ✓Agent: In transit, arriving Friday.User: Fine, then cancel it.            -> cancellation_flow  ✓                                          (but which order? gone)

Every individual routing decision is correct in isolation. The conversation is still broken, because a stateless function cannot hold a thread. Add memory to fix it and you have stopped being reactive.

Reactive: strengthsReactive: weaknesses
Millisecond latencyNo memory — cannot follow a conversation
Trivial to test exhaustivelyNo lookahead — cannot sequence actions
Cannot loop or thrashRule count explodes as cases multiply
Fully auditable: you can read every ruleSilent on anything unanticipated
Cheap — often no model call at allCannot improve without a human editing rules

Use reactive when the input space is small and known, the response is fixed, latency is critical, or the action is a safety reflex that must never wait on a planner.

Deliberative agents

A deliberative agent maintains an explicit model of the world, considers possible futures, and produces a plan — a sequence of actions expected to reach the goal — before acting.

Text
percepts ──▶ [ WORLD MODEL ]  what is true now                    │                    ▼             [ GOAL STATE ]   what should be true                    │                    ▼             [   PLANNER   ]  search over action sequences                    │                    ▼             [  PLAN: a1, a2, a3, ... ]                    │                    ▼             execute a1, a2, a3 ...

Worked example: booking a trip

Goal: be in Tokyo by 09:00 on 3 May, spend under 90,000 rupees, and do not travel overnight on a working day.

The agent's world model holds flight options, hotel availability, the user's calendar and the budget. The planner searches sequences:

Text
Candidate A: fly 2 May 14:00 (52,000)  + hotel 2 nights (24,000)             = 76,000, arrives 2 May 23:00, sleep, meeting 09:00  ✓Candidate B: fly 3 May 01:30 (38,000)  + hotel 1 night (12,000)             = 50,000, but departs during Fri night          ✗ violates                                                        constraint 3Candidate C: fly 1 May 09:00 (61,000)  + hotel 3 nights (36,000)             = 97,000                                   ✗ over budgetChosen: A. Total 76,000, 14,000 under budget, no overnight leg.

Only a deliberative agent can do this. Candidate B is cheaper by 26,000 and a reactive "pick the cheapest flight" rule would take it — and violate a constraint the user cares about. Evaluating a whole sequence against all constraints before committing is exactly the capability that planning buys.

Where it breaks

The world changes while you plan. This is the warehouse robot, and it is the fundamental problem. If your plan takes 3 seconds to compute and the environment changes every 1 second, your plans are always describing a world that has passed.

The model is wrong. A plan is only as good as the world model behind it. If the agent believes flight XY tickets are available and they sold out an hour ago, every step after that assumption is fiction. Deliberative agents fail confidently, because internal consistency feels like correctness.

Search cost grows brutally. With a branching factor of 5 and a depth of 3 you examine 125 sequences — fine. Depth 8 is 390,625. Depth 12 is over 244 million. Add one more option per step and depth 12 becomes 2.2 billion. Long-horizon planning without a good heuristic to prune with is not slow; it is impossible.

Deliberative: strengthsDeliberative: weaknesses
Handles multi-step, constrained goalsLatency measured in seconds
Can prove a plan satisfies constraintsRequires an accurate world model
Can compare alternatives before committingStale plans in dynamic environments
Explains itself — the plan is the explanationSearch cost explodes with horizon
Generalises to unseen goal combinationsBrittle when a mid-plan step fails

Use deliberative when actions have ordering dependencies, when constraints must be checked across the whole sequence, and when the environment is slow enough that a plan survives long enough to execute.

Hybrid architectures

The warehouse robot's problem has an obvious fix once stated: let the reflex interrupt the planner. That is the hybrid idea. Run a fast reactive layer that owns anything urgent, and a slow deliberative layer that owns anything that needs thought, with a rule for who wins.

Text
    percept       │       ├──────────────▶ REFLEX LAYER  (µs–ms)       │                obstacle? stop. fire? evacuate.       │                     │       │                     └── if triggered: ACT NOW, preempt       │       └──────────────▶ DELIBERATIVE LAYER  (100ms–s)                        update model, plan, refine intentions                             │                             └── if no reflex fired: ACT on plan

The layering rule that makes this work: the reactive layer always has priority, and the deliberative layer must be interruptible. If a plan cannot be abandoned mid-execution, you have two architectures stapled together, not a hybrid.

BDI: the standard way to organise the thinking layer

BDI stands for Belief–Desire–Intention, and it is the most widely used framing for hybrid agents. Three stores:

  • Beliefs — what the agent takes to be true about the world right now. Updated by perception. Beliefs can be wrong; that is the point of calling them beliefs rather than facts.
  • Desires — goals the agent would like to achieve. There can be many, and they can conflict ("resolve fast" and "never break policy" pull against each other).
  • Intentions — the subset of desires the agent has committed to and is actively pursuing. Commitment is the key idea: an intention persists across steps, so the agent does not abandon a goal the instant something else looks marginally more attractive.

Why commitment matters. Without it, an agent with desires "answer the customer" and "check for new tickets" ping-pongs — one step towards each, forever, finishing neither. Intentions give it the stubbornness to see something through. Too much stubbornness and it pursues a goal after the reason for it evaporated, so the deliberation cycle includes a reconsideration test.

A BDI support agent

Python
class BDIAgent:    def __init__(self):        self.beliefs = {}       # facts about the world        self.desires = []       # (goal, priority)        self.intention = None   # the one we are committed to    def perceive(self, event):        self.beliefs.update(event)    def deliberate(self):        # 1. Generate options consistent with current beliefs        options = [(g, p) for g, p in self.desires if self.feasible(g)]        # 2. Should we reconsider? Only if the current intention is        #    finished, impossible, or clearly outranked.        if self.intention:            still_ok = (self.feasible(self.intention)                        and not self.achieved(self.intention))            better = max(options, key=lambda o: o[1], default=(None, -1))            outranked = better[1] > self.priority_of(self.intention) + 1            if still_ok and not outranked:                return self.intention          # stay committed        # 3. Commit to the highest-priority feasible option        self.intention = max(options, key=lambda o: o[1])[0] if options \                         else None        return self.intention    def step(self, event):        self.perceive(event)        if reflex := self.check_reflexes():     # reactive layer wins            return reflex        goal = self.deliberate()        return self.next_action_for(goal)    def check_reflexes(self):        if self.beliefs.get("customer_says_legal"):            return "escalate_to_legal"          # never planned around        if self.beliefs.get("pii_detected"):            return "redact_and_halt"        return None

Read the outranked line carefully — better[1] > priority + 1, not > priority. That margin is the anti-dithering device. A competing goal must be meaningfully better, not marginally better, to steal commitment. Set the margin to zero and the agent flip-flops on noise; set it too high and it ignores genuine emergencies. This one constant is where a lot of BDI tuning lives.

And note that check_reflexes runs before deliberate and returns directly. A legal threat does not enter the planner as a high-priority desire to be weighed. It bypasses thinking entirely, which is exactly what you want for anything that must never be reasoned away.

Hybrid: strengthsHybrid: weaknesses
Fast where speed matters, thoughtful elsewhereTwo systems to build, test and reason about
Safety reflexes cannot be planned aroundLayer-interaction bugs are subtle and rare
Degrades gracefully — reflexes survive planner failureArbitration rules need real tuning
Commitment prevents goal thrashingHarder to explain a decision — which layer chose it?

Side by side

ReactiveDeliberativeHybrid / BDI
World modelNoneExplicit, centralBeliefs, possibly stale
MemoryNonePlan stateBeliefs + intentions
LookaheadZeroMulti-stepMulti-step, interruptible
Response timeMicroseconds100 ms – secondsReflex µs, plan ~seconds
Handles noveltyBadly — falls through rulesWell, if the model covers itWell
Dynamic environmentExcellentPoor — plans go staleGood
DebuggabilityTrivialModerate — read the planHard
Build effortHoursDaysWeeks
Compute per decisionNegligibleHighModerate to high
Typical useRouting, safety cut-offs, filtersTrip planning, research, schedulingRobotics, trading, live support

Choosing one

Four questions, answered in order. The first "yes" that forces an upgrade wins.

Text
1. Does the correct action depend on anything other than the   current input?       NO  ──▶ REACTIVE. Stop here. Do not build more.       YES ──▶ continue2. Do actions have to happen in a particular order, or must   constraints hold across a whole sequence?       NO  ──▶ REACTIVE + a memory store is usually enough       YES ──▶ continue3. Does the environment change faster than you can plan, OR is   there any action that must happen immediately and must never   be reasoned away?       NO  ──▶ DELIBERATIVE       YES ──▶ HYBRID4. (Sanity check) Can you afford to build, test and operate two   coordinated layers?       NO  ──▶ DELIBERATIVE + a hard-coded safety interlock outside               the agent. Simpler, and covers most of the benefit.
If your system...Build
Classifies or routes one input at a timeReactive
Answers questions from a fixed knowledge baseReactive with retrieval
Must satisfy a budget and a deadline togetherDeliberative
Researches a topic across many sourcesDeliberative
Controls something physicalHybrid, no exceptions
Trades, deploys, or moves moneyHybrid — the reflex layer is your circuit breaker
Holds a multi-turn conversation and takes actionsHybrid, lightweight

Choose the simplest architecture that can express the failure you are most afraid of. If your worst case is "wrong answer", deliberative is enough. If it is "acted too late", you need reflexes.

Where people get this wrong

Reaching for deliberation because the task sounds sophisticated. "Intelligent email triage" sounds like it needs planning. It is a classifier. Building a planner for it means paying seconds of latency and dollars of tokens to produce a one-word answer, with a new failure mode — the loop — that the classifier did not have.

Putting safety rules in the planner. A prompt saying "never issue a refund above 500 dollars" is a desire competing with other desires. Given enough pressure in a conversation, the model will construct a reason why this case is different. Safety rules belong in the reflex layer or, better, outside the agent in the service that executes the action, where no reasoning can reach them.

Building a hybrid where the layers never actually interact. A common shape: a reflex layer that checks three conditions which have never once fired in production, plus a planner doing all the work. That is a deliberative agent with extra code. Either the reflexes are load-bearing or delete them.

Letting the planner replan on every percept. This is BDI without intentions. The agent recomputes from scratch each step, picks a slightly different plan because the model is stochastic, and makes no progress on any of them. Commitment is not an optimisation; it is what makes multi-step behaviour possible.

What this means when you build one

Most real LLM systems end up as light hybrids, and they get there by accident and badly. You will build a deliberative agent because that is what the frameworks encourage, then discover it needs to never do a particular thing, and then bolt on checks. Do it deliberately instead.

Start by writing down the actions that must never be reasoned about — the escalations, the hard limits, the kill switches. Implement those as straight conditionals that run before the model call and return immediately. That is your reactive layer, and it should be small enough to read in one screen.

Then build the deliberative part for everything else, and make it interruptible: a step budget, a cancellation check between steps, and no operation so long that a reflex cannot preempt it.

If your agent holds a goal across multiple turns, add intentions — one field recording what it is currently committed to, and an explicit test for when to abandon it. Include the margin. An agent that switches goals only when something is clearly better looks far more competent than one that switches whenever something is slightly better, and the difference between those two behaviours is a single comparison operator.