Course Content
Applied AI Engineering: From Prompt to Production
9 sections · 29 lessons
Workflows or agents: choosing the right amount of autonomy
Once the agent worked, a teammate made a reasonable-sounding proposal: "Route everything through the agent. It can search, check balances and raise tickets, so it handles every case. One code path is simpler than five."
The team tested it instead of arguing. For one week, a copy of the traffic, about 6,000 questions, ran through both designs in the background. For plain policy questions, which were 70% of traffic, the agent was slower (6.8 seconds against 4.5), cost more than twice as much, and was slightly less accurate. It sometimes skipped the search and answered from general knowledge, and sometimes searched three times when once was enough. Generality has a price, and most questions did not need it.
This lesson gives you a way to decide how much autonomy each part of a system should have, and shows the design PolicyPal actually ships.
The autonomy ladder
The real question is: who decides the next step, your code or the model? Each rung up the ladder hands more of that decision to the model.
| Rung | Who decides the steps | PolicyPal example |
|---|---|---|
| Single call | Nobody; one step | Rewrite an answer more simply |
| Fixed chain | Code, always the same steps | Retrieve, answer, check grounding |
| Router plus chains | The model picks a path; code runs it | "Is this a policy question or a balance check?" |
| Bounded agent | The model picks each step, within budgets | Priya's two-part request |
| Open-ended agent | The model sets its own sub-goals over a long time | Not used in PolicyPal |
Each rung up gains flexibility and loses predictability. A fixed chain always takes the same path, so its cost, latency and failure modes are easy to know and test. An agent's path depends on the model's choices, so each run can differ, and testing needs many runs. Neither is better in general. The right rung is the lowest one that handles the task.
Same traffic, two designs
Here is the week of shadow traffic, split by the kind of question.
| Kind of question | Share | Workflow: success, latency, cost | Agent: success, latency, cost |
|---|---|---|---|
| Plain policy question | 70% | 91%, 4.5 s, $0.0085 | 89%, 6.8 s, $0.019 |
| Own leave balance | 9% | 98%, 1.8 s, $0.004 | 97%, 4.1 s, $0.012 |
| IT ticket request | 7% | 95%, 2.6 s, $0.005 | 94%, 5.0 s, $0.014 |
| Several intents in one message | 7% | 61%, 4.6 s, $0.009 | 90%, 7.2 s, $0.021 |
The pattern is clear. Where the path is predictable, the workflow wins on every number. Where the message mixes intents, the fixed workflow fails badly: asked about leave and a VPN problem together, the policy chain answers the leave question and silently drops the VPN request. That is exactly where the agent earns its cost.
Where agents earn their cost
Use an agent when the next step genuinely depends on what the previous step returned, when requests combine several intents in ways you cannot list in advance, and when a slower, costlier, less predictable answer is still acceptable.
Prefer a workflow when you can write down the steps, when latency or cost limits are tight, when you need the same behaviour every time for audit or compliance, and when an action is high-stakes. It is not a judgement on how clever the model is. It is a judgement on how much variation the task can tolerate.
The choice is not made once. The agent's traces are a record of the paths real requests need. If you read them monthly, you will find paths that repeat. At Harbourline, "check my balance, then explain whether I can take next week off" appeared hundreds of times with the same two steps. A repeated path is a workflow waiting to be written: turn it into a fixed chain, give it a route, and the agent handles a smaller, stranger remainder. Moving the other way also happens. If a workflow keeps growing special cases for combinations it cannot foresee, that is a sign the task needs the model to choose the steps.
PolicyPal's design: a router in front
PolicyPal puts a router in front of everything. It classifies each message into one of six routes, and each route has its own path.
1# policypal/dispatch.py2from policypal.agent import run_agent3from policypal.flows import balance_flow, handover, out_of_scope, policy_flow, ticket_flow45ROUTES = {6 "policy_question": policy_flow, # retrieve, answer, check grounding7 "leave_balance": balance_flow, # one tool call, then a templated answer with the policy8 "it_ticket": ticket_flow, # fill ticket fields, show the confirmation card9 "sensitive_hr": handover, # no model answer; open an HR partner case10 "out_of_scope": out_of_scope, # a fixed, polite reply naming the right channel11}1213def handle(message: str, user, session) -> dict:14 route = session.router.classify(message) # one of six labels15 session.trace.set_attribute("route", route)16 if route == "multi_step":17 return run_agent(session.llm, session.agent_system, session.history(message), user)18 return ROUTES[route](message, user, session)Most routes are workflows. sensitive_hr deliberately uses no model answer at all: a question about harassment or a health condition goes straight to a trained person, with the conversation attached. Only multi_step reaches the agent. The router is the single place that decides autonomy, and it is easy to measure: for each message, was the route right?
For now the router is a prompted call to the mid-size model with structured output limited to the six labels. It is 94.8% accurate on a labelled set, but it adds about 850 milliseconds and $0.001 to every question. Section 5 replaces it with a small fine-tuned model that does the same job in about 15 milliseconds.
The failure you most need to watch is a multi-intent message routed to a single workflow, because that silently drops part of the request. PolicyPal's router instructions say "if the message asks for two different things, choose multi_step", and the eval set in Section 6 has a group of two-intent messages just to measure this.
Check your understanding
0 of 3 answered
1.For plain policy questions, the agent was slightly less accurate than the fixed workflow. What is the most likely reason?
2.Why does the sensitive_hr route use no model-generated answer at all?
3.A user writes, "What's the policy on sabbaticals, and can you reset my email password?" The router labels it policy_question. What happens, and what is the fix?