Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Scenario – 4: Unstable Conditional Routing
What you need to know
The scenario: the same kind of question goes to the billing node one time and the technical node the next. The router is an LLM call inside a conditional edge.
Why routing wobbles
Asking a model "which team should handle this?" and parsing prose ("I think this is probably billing") fails in two ways: the wording varies, and sampling makes the choice itself vary. Both are fixable.
Make the decision constrained
1from typing import Literal2from pydantic import BaseModel34class Route(BaseModel):5 destination: Literal["billing", "technical", "general"]6 confidence: float78router_llm = llm.with_structured_output(Route) # llm created with temperature=0910RULES = {"refund": "billing", "invoice": "billing", "error code": "technical"}1112def route(state) -> Literal["billing", "technical", "general"]:13 text = state["question"].lower()14 for word, dest in RULES.items():15 if word in text:16 return dest # obvious cases never reach the LLM17 choice = router_llm.invoke(state["question"])18 if choice.confidence < 0.6:19 return "general" # safe default, which can ask a clarifying question20 return choice.destination2122builder.add_conditional_edges("classify", route,23 {"billing": "billing", "technical": "technical", "general": "general"})The explicit path map keeps the graph inspectable, and the Literal return type documents every exit.
Measure it
| Signal | What it tells you |
|---|---|
| Accuracy on 200–300 labelled real queries | Whether routing is good enough (aim high, e.g. 95%+) |
| Confusion between two routes | The taxonomy may be wrong; merge them or add a clarifying node |
| One route dominating | A prompt bias, often toward the first option listed |
| Low-confidence rate | How often users get the default path |
Log every decision with its confidence. Routing becomes a tracked metric, and a prompt change that drops accuracy is blocked in CI.
A real-life example
Scenario, numbers made up. A broadband provider's support graph routes chats to billing, technical or general. The LLM router is asked for a one-word answer and the code checks if "billing" in reply. Replies like "Not billing — technical" go to billing. Routing accuracy on 250 labelled chats is 81%, and running the same chat five times gives two different routes in 12% of cases.
The team adds keyword rules (they catch 55% of chats), an enum structured output at temperature 0, and a default to general below 0.6 confidence, where the node asks one clarifying question. Accuracy rises to 96%, repeat-run disagreement falls below 1%, and misrouted-ticket transfers drop by about two-thirds.
Follow-up questions to expect
- "Why not fine-tune a classifier?" — Once you have a few thousand labelled routing decisions, a small trained classifier is cheaper and more stable; the LLM router is a good start and a label generator.
- "Does temperature 0 make it deterministic?" — Much more stable, but not guaranteed identical on every provider. The schema and validation handle the rest.
- "What if a question needs two teams?" — Allow a list of destinations and fan out with
Send, or route to a node that handles multi-topic questions.