Applied AI Engineering: From Prompt to Production

Course Content

Applied AI Engineering: From Prompt to Production

9 sections · 29 lessons

Workflows or agents: choosing the right amount of autonomy


Once the agent worked, a teammate made a reasonable-sounding proposal: "Route everything through the agent. It can search, check balances and raise tickets, so it handles every case. One code path is simpler than five."

The team tested it instead of arguing. For one week, a copy of the traffic, about 6,000 questions, ran through both designs in the background. For plain policy questions, which were 70% of traffic, the agent was slower (6.8 seconds against 4.5), cost more than twice as much, and was slightly less accurate. It sometimes skipped the search and answered from general knowledge, and sometimes searched three times when once was enough. Generality has a price, and most questions did not need it.

This lesson gives you a way to decide how much autonomy each part of a system should have, and shows the design PolicyPal actually ships.

Where each message goesRouterPolicy question:fixed chain, 70%Leave balance: one tool, 9%IT ticket:confirmation card, 7%Two intents:bounded agent, 7%Sensitive HR: a person, 3%Out of scope:fixed reply, 4%
93% of traffic takes a fixed path; the agent is kept for the 7% whose steps cannot be known in advance.

The autonomy ladder

The real question is: who decides the next step, your code or the model? Each rung up the ladder hands more of that decision to the model.

RungWho decides the stepsPolicyPal example
Single callNobody; one stepRewrite an answer more simply
Fixed chainCode, always the same stepsRetrieve, answer, check grounding
Router plus chainsThe model picks a path; code runs it"Is this a policy question or a balance check?"
Bounded agentThe model picks each step, within budgetsPriya's two-part request
Open-ended agentThe model sets its own sub-goals over a long timeNot used in PolicyPal

Each rung up gains flexibility and loses predictability. A fixed chain always takes the same path, so its cost, latency and failure modes are easy to know and test. An agent's path depends on the model's choices, so each run can differ, and testing needs many runs. Neither is better in general. The right rung is the lowest one that handles the task.

Same traffic, two designs

Here is the week of shadow traffic, split by the kind of question.

Kind of questionShareWorkflow: success, latency, costAgent: success, latency, cost
Plain policy question70%91%, 4.5 s, $0.008589%, 6.8 s, $0.019
Own leave balance9%98%, 1.8 s, $0.00497%, 4.1 s, $0.012
IT ticket request7%95%, 2.6 s, $0.00594%, 5.0 s, $0.014
Several intents in one message7%61%, 4.6 s, $0.00990%, 7.2 s, $0.021

The pattern is clear. Where the path is predictable, the workflow wins on every number. Where the message mixes intents, the fixed workflow fails badly: asked about leave and a VPN problem together, the policy chain answers the leave question and silently drops the VPN request. That is exactly where the agent earns its cost.

Where agents earn their cost

Use an agent when the next step genuinely depends on what the previous step returned, when requests combine several intents in ways you cannot list in advance, and when a slower, costlier, less predictable answer is still acceptable.

Prefer a workflow when you can write down the steps, when latency or cost limits are tight, when you need the same behaviour every time for audit or compliance, and when an action is high-stakes. It is not a judgement on how clever the model is. It is a judgement on how much variation the task can tolerate.

The choice is not made once. The agent's traces are a record of the paths real requests need. If you read them monthly, you will find paths that repeat. At Harbourline, "check my balance, then explain whether I can take next week off" appeared hundreds of times with the same two steps. A repeated path is a workflow waiting to be written: turn it into a fixed chain, give it a route, and the agent handles a smaller, stranger remainder. Moving the other way also happens. If a workflow keeps growing special cases for combinations it cannot foresee, that is a sign the task needs the model to choose the steps.

PolicyPal's design: a router in front

PolicyPal puts a router in front of everything. It classifies each message into one of six routes, and each route has its own path.

Python
# policypal/dispatch.pyfrom policypal.agent import run_agentfrom policypal.flows import balance_flow, handover, out_of_scope, policy_flow, ticket_flowROUTES = {    "policy_question": policy_flow,    # retrieve, answer, check grounding    "leave_balance": balance_flow,     # one tool call, then a templated answer with the policy    "it_ticket": ticket_flow,          # fill ticket fields, show the confirmation card    "sensitive_hr": handover,          # no model answer; open an HR partner case    "out_of_scope": out_of_scope,      # a fixed, polite reply naming the right channel}def handle(message: str, user, session) -> dict:    route = session.router.classify(message)          # one of six labels    session.trace.set_attribute("route", route)    if route == "multi_step":        return run_agent(session.llm, session.agent_system, session.history(message), user)    return ROUTES[route](message, user, session)

Most routes are workflows. sensitive_hr deliberately uses no model answer at all: a question about harassment or a health condition goes straight to a trained person, with the conversation attached. Only multi_step reaches the agent. The router is the single place that decides autonomy, and it is easy to measure: for each message, was the route right?

For now the router is a prompted call to the mid-size model with structured output limited to the six labels. It is 94.8% accurate on a labelled set, but it adds about 850 milliseconds and $0.001 to every question. Section 5 replaces it with a small fine-tuned model that does the same job in about 15 milliseconds.

The failure you most need to watch is a multi-intent message routed to a single workflow, because that silently drops part of the request. PolicyPal's router instructions say "if the message asks for two different things, choose multi_step", and the eval set in Section 6 has a group of two-intent messages just to measure this.

Check your understanding

0 of 3 answered

1.For plain policy questions, the agent was slightly less accurate than the fixed workflow. What is the most likely reason?

2.Why does the sensitive_hr route use no model-generated answer at all?

3.A user writes, "What's the policy on sabbaticals, and can you reset my email password?" The router labels it policy_question. What happens, and what is the fix?