Building AI Features in Python Backends

Routing requests with a classifier


A classifier by itself does nothing for ShipFast. Its value comes from what happens next: the damaged-parcel message reaching the claims team in minutes, the reschedule landing in a queue where a tool can book the new slot. That step is routing, and it is where many teams make a quiet mistake: they let the model choose the queue.

It seems natural. Give the model the list of queues and let it pick one. But queues are an internal detail that changes: teams merge, a new claims desk opens, business customers get a dedicated queue. Every change would mean editing a prompt and re-running an evaluation. Worse, the model has no idea which queue is overloaded, which customer is a business account or whether a reschedule has enough information to be booked automatically. Your code knows all of that.

So ShipFast splits the job: the model labels, the code routes.

Share of messages agents move out of each queue1.72.81.37.95.701234reschedulingclaimsescalationsgeneralPercent rerouted per day. Escalations lose most messages to rescheduling.
Reroutes out of escalations turned out to be a label-definition question, which only shows up because code, not the model, owns routing.

Rules first, model second

Before any model call, a regular expression catches the most common message that needs no understanding at all: "where is SF12345678?".

Python
# shipfast/routing.pyimport refrom shipfast.schemas import Extraction, Intent, QueueTRACKING_ONLY = re.compile(    r"\s*(where is|status( of)?|track)?\s*(my )?(parcel|order)?\s*[:#-]?\s*SF\d{8}\s*\??\s*",    re.IGNORECASE)def is_status_question(text: str) -> bool:    return TRACKING_ONLY.fullmatch(text) is not None

It matches "where is SF12345678?", "track SF12345678" and a bare "SF12345678", and it does not match "where is SF12345678, deliver tomorrow", because fullmatch requires the whole message to be a status question. That is about 12% of ShipFast's traffic. Those messages go to the status bot, which reads the parcel's tracking record and replies instantly.

The saving is real: 12% of 40,000 is 4,800 messages a day that skip two model calls each. At ShipFast's prices, that is about $36 a day, over $1,000 a month, for a ten-line function that is never wrong. The status bot also answers in 200 ms instead of 2 seconds.

From label to queue

Python
# shipfast/routing.py (continued)ROUTES: dict[Intent, Queue] = {    Intent.RESCHEDULE: Queue.RESCHEDULING,    Intent.ADDRESS_CHANGE: Queue.ADDRESS_DESK,    Intent.DAMAGED_PARCEL: Queue.CLAIMS,    Intent.COMPLAINT: Queue.ESCALATIONS,    Intent.OTHER: Queue.GENERAL,    Intent.UNKNOWN: Queue.HUMAN_TRIAGE,}def route(intent: Intent, fields: Extraction | None, is_business: bool) -> tuple[Queue, str]:    queue = ROUTES[intent]    # A reschedule without a date, or an address change without a PIN code,    # cannot be actioned by the desk automation: a person must ask.    if intent is Intent.RESCHEDULE and (fields is None or fields.delivery_date is None):        queue = Queue.HUMAN_TRIAGE    if intent is Intent.ADDRESS_CHANGE and (fields is None or fields.pincode is None):        queue = Queue.HUMAN_TRIAGE    high = intent is Intent.DAMAGED_PARCEL or (intent is Intent.COMPLAINT and is_business)    return queue, "high" if high else "normal"

Three kinds of rule live here, and none of them belong in a prompt.

The mapping. A dictionary from intent to queue. When ShipFast opens a new desk, this dictionary changes and nothing else does. Because Intent is an enum and the dictionary covers every member, a test can check that no intent is left without a queue.

Completeness checks. The rescheduling queue is worked by a tool that books slots automatically. A reschedule without a date cannot be booked, so it goes to a person who will ask the customer. This is where extraction's "null over guessing" pays off: a missing date is visible, so it is routed correctly.

Priority. Damaged parcels are always high priority, because claims have deadlines. Complaints from business customers are high too. The model does not know who is a business customer, and should not.

Measuring routing quality

Your evaluation set tells you how often the label is right. Production tells you something better: how often agents move a message to a different queue. ShipFast logs a rerouted event whenever an agent changes a message's queue, with the old and new queue and the prompt version.

QueueMessages per dayRerouted by agentsRate
rescheduling11,2001901.7%
address_desk3,9001102.8%
claims3,100401.3%
escalations5,8004607.9%
general4,9002805.7%
human_triage2,400n/an/a

Escalations stand out at 7.9%. Reading the rerouted messages showed most were late-delivery complaints that also asked for a new date, which agents moved to rescheduling. That points at a product decision (should a complaint with a date request go to rescheduling?) rather than a model error. Routing metrics often tell you about your label definitions and your process, not only about the model.

Routing to the next step, not only to a queue

The label can also decide what the service does next. In ShipFast, only reschedule and address_change trigger extraction, because only they need fields; that saves an extraction call on more than half of all messages. The label could choose a model too: damaged-parcel drafts, which need care, could use a stronger model, while simple reschedule confirmations might not need a model at all. A complete reschedule with a date and a window can be confirmed with a template, which is free, instant and never says anything unexpected. Section 4's project leaves this as an exercise, because it is a good example of the best model call being the one you do not make.

Check your understanding

0 of 3 answered

1.Why does ShipFast use fullmatch for the status-question regex instead of search?

2.A reschedule message is classified correctly, but extraction returns no delivery date. Where should it go?

3.Agents reroute 7.9% of messages out of escalations, mostly to rescheduling. What is the best first interpretation?