Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Scenario – 4: Unstable Conditional Routing


What you need to know

The scenario: the same kind of question goes to the billing node one time and the technical node the next. The router is an LLM call inside a conditional edge.

Why routing wobbles

Asking a model "which team should handle this?" and parsing prose ("I think this is probably billing") fails in two ways: the wording varies, and sampling makes the choice itself vary. Both are fixable.

Make the decision constrained

Python
from typing import Literalfrom pydantic import BaseModelclass Route(BaseModel):    destination: Literal["billing", "technical", "general"]    confidence: floatrouter_llm = llm.with_structured_output(Route)    # llm created with temperature=0RULES = {"refund": "billing", "invoice": "billing", "error code": "technical"}def route(state) -> Literal["billing", "technical", "general"]:    text = state["question"].lower()    for word, dest in RULES.items():        if word in text:            return dest                              # obvious cases never reach the LLM    choice = router_llm.invoke(state["question"])    if choice.confidence < 0.6:        return "general"                             # safe default, which can ask a clarifying question    return choice.destinationbuilder.add_conditional_edges("classify", route,    {"billing": "billing", "technical": "technical", "general": "general"})

The explicit path map keeps the graph inspectable, and the Literal return type documents every exit.

Measure it

SignalWhat it tells you
Accuracy on 200–300 labelled real queriesWhether routing is good enough (aim high, e.g. 95%+)
Confusion between two routesThe taxonomy may be wrong; merge them or add a clarifying node
One route dominatingA prompt bias, often toward the first option listed
Low-confidence rateHow often users get the default path

Log every decision with its confidence. Routing becomes a tracked metric, and a prompt change that drops accuracy is blocked in CI.

A real-life example

Scenario, numbers made up. A broadband provider's support graph routes chats to billing, technical or general. The LLM router is asked for a one-word answer and the code checks if "billing" in reply. Replies like "Not billing — technical" go to billing. Routing accuracy on 250 labelled chats is 81%, and running the same chat five times gives two different routes in 12% of cases.

The team adds keyword rules (they catch 55% of chats), an enum structured output at temperature 0, and a default to general below 0.6 confidence, where the node asks one clarifying question. Accuracy rises to 96%, repeat-run disagreement falls below 1%, and misrouted-ticket transfers drop by about two-thirds.

Follow-up questions to expect

  • "Why not fine-tune a classifier?" — Once you have a few thousand labelled routing decisions, a small trained classifier is cheaper and more stable; the LLM router is a good start and a label generator.
  • "Does temperature 0 make it deterministic?" — Much more stable, but not guaranteed identical on every provider. The schema and validation handle the rest.
  • "What if a question needs two teams?" — Allow a list of destinations and fan out with Send, or route to a node that handles multi-topic questions.