AI Product Engineering: Shipping LLM Features That Last

Course Content

AI Product Engineering: Shipping LLM Features That Last

6 sections · 22 lessons

Choosing a model: quality, cost and latency together


TiffinGo chose its model last, on purpose. Until the contract, eval set, retrieval and tools were in place, any comparison would have measured the wrong thing: models failing on missing data rather than on the task. Now the question is fair: given everything around it, which model should run this feature?

Public leaderboards do not answer that. They measure other tasks, in other languages, with other prompts. A model that tops a coding benchmark may be no better at reading "2 roti kam aaye". The only measurement that counts is your eval set, run through your full pipeline.

Quality alone is also not enough. A model that is 1 point better and costs three times as much is only better if that 1 point is worth the money. This lesson puts all three numbers on the same table.

Three tiers through the same pipelineLarge116 of 120₹1,87010.2 s8 of 8Medium114 of 120₹7455.1 s8 of 8Small103 of 120₹3702.6 s7 of 8TierEval passCost a dayp95Food safetyDrafts are made when the ticket arrives, so latency is mostly hidden.
Twenty more right drafts a day were worth about ₹460 and cost about ₹1,125 more, so the middle tier won on the lesson's own arithmetic.

Run the candidates through the gate

The team ran v6 on three tiers of model from one provider, which we will call large, medium and small. Each ran the full eval three times, through the same retrieval, tools, validator and pricing code.

Model tierEval passFood safetyp50 latencyp95 latencyCost per ticketCost per day
Large116 of 1208 of 84.6 s10.2 s₹1.56₹1,870
Medium114 of 1208 of 82.4 s5.1 s₹0.62₹745
Small103 of 1207 of 81.2 s2.6 s₹0.31₹370

The small model is out at once: it fails a food-safety row, and the gate does not allow that for any price. The real choice is between large and medium.

Cost per ticket, from your own token counts

Cost comes from your measured token usage, not the provider's example. TiffinGo's drafts average about 3,600 input tokens (a 1,500-token system prompt, 600 for the ticket and order, 900 for policy sections, and about 600 more on tickets that make a tool call) and 260 output tokens.

Python
USD_TO_INR = 88.0def cost_per_ticket_inr(uncached_in: int, cached_in: int, out: int,                        price_in: float, price_out: float, cache_factor: float = 0.1) -> float:    """Prices are US dollars per million tokens; cached input is billed at cache_factor."""    usd = (uncached_in * price_in           + cached_in * price_in * cache_factor           + out * price_out) / 1_000_000    return usd * USD_TO_INRfor tier, price_in, price_out in [("large", 5, 25), ("medium", 2, 10), ("small", 1, 5)]:    print(tier, round(cost_per_ticket_inr(2100, 1500, 260, price_in, price_out), 2))# large 1.56   medium 0.62   small 0.31

The prices are illustrative per-million-token rates for three tiers; use your provider's current price list. Caching the stable 1,500-token system prompt matters: for the medium tier it cuts the cost from ₹0.86 to ₹0.62 per ticket, about 28%, for no change in quality. It is the cheapest cost fix available, which is one more reason to keep per-ticket data out of the system prompt. Some providers charge a small premium to write the cache; at 1,200 tickets a day it is negligible.

Put a price on the quality difference

The large model gets 2 more tickets of 120 right: about 1.7%, or about 20 more correct drafts a day at 1,200 tickets. What is a wrong draft worth?

From section 2: an agent catches about 85% of wrong drafts, and fixing one takes about 3 minutes, around ₹12.50 of agent time. The other 15% slip through and cost about ₹80 on average. So one wrong draft costs about 0.85 × ₹12.50 + 0.15 × ₹80, roughly ₹23.

Twenty fewer wrong drafts a day saves about ₹460. The large model costs about ₹1,125 a day more. Medium wins, by a clear margin. For the small model the same arithmetic runs the other way: about 110 more wrong drafts a day, costing around ₹2,500, to save ₹375, and that is before its food-safety miss.

This arithmetic will not always favour the middle. If TiffinGo's wrong drafts were expensive, for example in a pharmacy app where a wrong answer is a health risk, 1.7% might be worth far more than ₹1,125 a day.

How much does latency matter here?

The instinct is to prefer the fastest model. First ask when the user actually waits.

TiffinGo generates the draft when the ticket is created, not when the agent opens it. The median ticket waits four minutes in the queue. A draft that takes 2.4 or even 10 seconds is ready long before anyone looks. Latency is almost invisible to agents, so it barely enters the decision.

The same model in a live chat, where a customer watches a typing indicator, would face a very different constraint: a p95 of 10 seconds would feel broken. Always ask where the user's wait actually is before paying for speed. And when latency does matter, look at p95, not the average; the slow tail is what people remember.

The choice is not permanent

Record the decision and its numbers in the changelog, and re-run the comparison when something changes: a new model generation, a price change, a large change in volume, or a new product need like live chat. Always pin an exact model id in the release config. When a provider updates a model behind an unpinned name, your eval results silently stop describing what is running.

TiffinGo re-runs the comparison every quarter. It costs about ₹800 in model calls and one afternoon.

Check your understanding

0 of 3 answered

1.The large model is 1.7 points better and costs ₹1,125 a day more. A wrong draft costs about ₹23. Which is right at 1,200 tickets a day?

2.Why does latency barely matter in TiffinGo's model choice?

3.Why pin an exact model id in the release config?