Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

How would you choose the LLM for a customer-support assistant handling 2 million calls a month?


Where the support bill actually fallsEvery ticket on the frontier model• About 18,000 dollars a month at 2M calls• Quality ceiling on easy and hard alike• No data on which tickets are easy• One provider, one point of failureRouted by intent• About 70% of tickets to a small model• Refunds and warrantystay on the large one• Blended bill near 5,000 dollars• Eval set checks every model swap
Routing, not picking one model, is where the saving comes from, and the eval set is what makes the small model safe to use.

What you need to know

The interviewer is not asking which model is best. They want to see that you turn a vague choice into a measured one. The honest answer is "I don't know yet, and here is how I'd find out cheaply".

The four constraints that decide it

ConstraintQuestion to askHow it narrows the choice
Quality barWhat does a wrong answer cost, and who pays?A wrong refund answer costs money; a wrong FAQ answer costs a re-ask
LatencyWhat p95 does the chat UI tolerate?A 2-second budget rules out slow reasoning modes for most turns
Data residencyCan ticket text leave the country or the network?May force a regional endpoint or self-hosting
Cost ceilingWhat can we spend per month, or per ticket?Sets how much traffic can go to an expensive model

Why start on a strong hosted model

At launch you do not know which tickets are hard. Running real traffic through the strongest model you can afford gives you a quality ceiling and a log of real inputs. Put the call behind a thin wrapper so the model name is one config value. Changing models later is then a config change, not a code change.

Put numbers on the cost

Assume each call sends about 1,500 input tokens and returns 300 output tokens. With illustrative prices of $3 per million input tokens and $15 per million output tokens for a frontier model:

Text
input:  2,000,000 calls x 1,500 tokens = 3.0B tokens  -> $9,000output: 2,000,000 calls x   300 tokens = 0.6B tokens  -> $9,000total ≈ $18,000 per month on the frontier model alone

A small model priced at a twentieth of that costs under $1,000 for the same traffic. If 75% of tickets are easy enough for it, the blended bill falls to roughly $5,000. That is why routing, not model choice alone, is where the savings come from.

The plan

  1. Launch — frontier model behind a wrapper; log model, tokens, latency and cost per request.
  2. Build the eval — 200 real tickets with labelled correct answers, covering refunds, order status, account issues and edge cases.
  3. Compare tiers — run a frontier, a mid-tier and a small open-weights model (Llama, Qwen or Mistral class, served with vLLM) over the eval set.
  4. Route — send the easy intents to the small model; escalate low-confidence or high-risk tickets to the big one.
Python
MODELS = {"small": "small-instruct", "large": "frontier-model"}   # config, not codedef answer(ticket):    tier = "small" if router.is_easy(ticket) else "large"    reply = llm.complete(model=MODELS[tier], prompt=build_prompt(ticket))    log_usage(ticket.id, tier, reply.usage, reply.latency_ms)    return reply

The router can be a small classifier trained on ticket intent, or a rule such as "refunds and legal questions always go large". The log line is what lets you compute cost per resolved ticket later.

When self-hosting wins

Self-hosting pays off only when you keep GPUs busy. A rented GPU costs the same per hour whether it serves 10 requests or 10,000. As a rough rule, if you cannot keep a node well above half utilised most of the day, per-token API pricing is cheaper. The exception is a hard data-residency or air-gap rule. Then self-hosting is the requirement, not an optimisation.

A real-life example

Scenario (illustrative numbers). An Indian electronics marketplace runs a support assistant on a frontier model. The bill is about ₹15 lakh a month. The team logs every call for three weeks and labels 200 tickets.

They find that 72% of tickets are order status, delivery date and return-window questions. A small open-weights model answers 94% of those correctly on the eval, against 97% for the frontier model. On refund disputes and warranty claims the small model drops to 71%, so those stay on the large model.

They route by intent. The small model now takes about 70% of traffic, the bill falls to roughly ₹5 lakh, and the escalation-to-human rate does not move. They keep the eval set in CI so any model swap is checked the same way.

Follow-up questions to expect

  • "Why not start with the small model to save money?" — You don't yet know where it fails. Starting strong gives you a quality ceiling and real data to measure the small model against.
  • "How do you set the routing threshold?" — On the eval set: pick the threshold where the small model's accuracy on routed tickets meets the bar, then watch the escalation rate in production.
  • "What if the provider has an outage?" — The wrapper holds a fallback model from a second provider, tested on the same eval, and a circuit breaker switches to it.