Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
How would you choose the LLM for a customer-support assistant handling 2 million calls a month?
What you need to know
The interviewer is not asking which model is best. They want to see that you turn a vague choice into a measured one. The honest answer is "I don't know yet, and here is how I'd find out cheaply".
The four constraints that decide it
| Constraint | Question to ask | How it narrows the choice |
|---|---|---|
| Quality bar | What does a wrong answer cost, and who pays? | A wrong refund answer costs money; a wrong FAQ answer costs a re-ask |
| Latency | What p95 does the chat UI tolerate? | A 2-second budget rules out slow reasoning modes for most turns |
| Data residency | Can ticket text leave the country or the network? | May force a regional endpoint or self-hosting |
| Cost ceiling | What can we spend per month, or per ticket? | Sets how much traffic can go to an expensive model |
Why start on a strong hosted model
At launch you do not know which tickets are hard. Running real traffic through the strongest model you can afford gives you a quality ceiling and a log of real inputs. Put the call behind a thin wrapper so the model name is one config value. Changing models later is then a config change, not a code change.
Put numbers on the cost
Assume each call sends about 1,500 input tokens and returns 300 output tokens. With illustrative prices of $3 per million input tokens and $15 per million output tokens for a frontier model:
input: 2,000,000 calls x 1,500 tokens = 3.0B tokens -> $9,000output: 2,000,000 calls x 300 tokens = 0.6B tokens -> $9,000total ≈ $18,000 per month on the frontier model aloneA small model priced at a twentieth of that costs under $1,000 for the same traffic. If 75% of tickets are easy enough for it, the blended bill falls to roughly $5,000. That is why routing, not model choice alone, is where the savings come from.
The plan
- Launch — frontier model behind a wrapper; log model, tokens, latency and cost per request.
- Build the eval — 200 real tickets with labelled correct answers, covering refunds, order status, account issues and edge cases.
- Compare tiers — run a frontier, a mid-tier and a small open-weights model (Llama, Qwen or Mistral class, served with vLLM) over the eval set.
- Route — send the easy intents to the small model; escalate low-confidence or high-risk tickets to the big one.
1MODELS = {"small": "small-instruct", "large": "frontier-model"} # config, not code23def answer(ticket):4 tier = "small" if router.is_easy(ticket) else "large"5 reply = llm.complete(model=MODELS[tier], prompt=build_prompt(ticket))6 log_usage(ticket.id, tier, reply.usage, reply.latency_ms)7 return replyThe router can be a small classifier trained on ticket intent, or a rule such as "refunds and legal questions always go large". The log line is what lets you compute cost per resolved ticket later.
When self-hosting wins
Self-hosting pays off only when you keep GPUs busy. A rented GPU costs the same per hour whether it serves 10 requests or 10,000. As a rough rule, if you cannot keep a node well above half utilised most of the day, per-token API pricing is cheaper. The exception is a hard data-residency or air-gap rule. Then self-hosting is the requirement, not an optimisation.
A real-life example
Scenario (illustrative numbers). An Indian electronics marketplace runs a support assistant on a frontier model. The bill is about ₹15 lakh a month. The team logs every call for three weeks and labels 200 tickets.
They find that 72% of tickets are order status, delivery date and return-window questions. A small open-weights model answers 94% of those correctly on the eval, against 97% for the frontier model. On refund disputes and warranty claims the small model drops to 71%, so those stay on the large model.
They route by intent. The small model now takes about 70% of traffic, the bill falls to roughly ₹5 lakh, and the escalation-to-human rate does not move. They keep the eval set in CI so any model swap is checked the same way.
Follow-up questions to expect
- "Why not start with the small model to save money?" — You don't yet know where it fails. Starting strong gives you a quality ceiling and real data to measure the small model against.
- "How do you set the routing threshold?" — On the eval set: pick the threshold where the small model's accuracy on routed tickets meets the bar, then watch the escalation rate in production.
- "What if the provider has an outage?" — The wrapper holds a fallback model from a second provider, tested on the same eval, and a circuit breaker switches to it.