Course Content
Enterprise AI Solutions Architecture
13 sections · 29 lessons
Unit Economics of an AI Feature
In the pilot, Meridian's finance partner asked a straightforward question: "What does one policy conversation cost?" The engineer answered, "About two cents a call." The first monthly invoice was three times the estimate. Nobody had lied. The estimate counted the main answer call and nothing else.
A single staff question at Meridian involves a guard check, a query rewrite, an embedding, a rerank, the answer itself, and sometimes a retry or a fallback. A conversation carries its history into later turns, so each turn costs more than the last. Letters have a check pass and are sometimes regenerated. Evaluation runs and shadow tests use real tokens too. And none of that includes the GPU, the search service, the logging, or the people who keep it all honest.
This lesson builds Meridian's unit economics properly. It also shows something that surprises most engineers: at this scale, tokens are one of the smallest costs in the system.
Count every call in a unit of work
Pick units that match how the business thinks: a conversation, a summary, a letter. Then list every call a unit makes. Meridian's staff ask an average of 2.2 questions per conversation, and each later turn carries about 800 tokens of history.
1PRICE_IN, PRICE_OUT = 3.00 / 1e6, 15.00 / 1e6 # illustrative dollars per token23def call_cost(tokens_in: int, tokens_out: int) -> float:4 return tokens_in * PRICE_IN + tokens_out * PRICE_OUT56# 2.2 questions: turn one, then 1.2 later turns carrying history, plus a rerank per question7conversation = call_cost(5_500, 350) + 1.2 * call_cost(6_300, 350) + 2.2 * 0.0018summary = 1.05 * call_cost(9_000, 500) # 5% regenerated9letter = 1.4 * (call_cost(7_000, 900) + call_cost(8_000, 200)) # draft plus check, 1.4 drafts1011per_day = (3_000 / 2.2) * conversation + 600 * summary + 70 * letter12per_year = per_day * 250 * 1.10 # plus 10% for evaluation and shadow runs13print(f"conversation ${conversation:.3f} summary ${summary:.3f} letter ${letter:.3f}")14print(f"tokens per day ${per_day:,.0f} per year ${per_year:,.0f}")The model uses the token budgets from section 4 and the illustrative prices from section 5. The multipliers carry the hidden calls: 1.2 later turns per conversation, 5% of summaries regenerated after a must-include failure, 1.4 drafts per letter because specialists sometimes change the plan, and 10% extra for evaluation and shadow traffic. The guard and rewrite calls do not appear because they run on the self-hosted GPU, which is a fixed cost.
It prints a conversation at about 5.3 cents, a summary at 3.6 cents and a letter at 8.6 cents. Meridian's 600 daily summaries are the 70 precomputed for new hardship cases plus about 530 that staff request on other cases with arrears. Across all traffic, tokens come to about $100 a day, or roughly $27,500 a year.
Variable, fixed and people costs
Tokens are the variable cost: they rise and fall with use. Two other kinds of cost sit beside them.
| Cost | Kind | Per year |
|---|---|---|
| Tokens, including evaluation and shadow runs | Variable | $27,500 |
| Self-hosted guard and rewrite GPU | Fixed | $24,000 |
| Search index service | Fixed | $14,400 |
| Tracing, logs and audit storage | Fixed | $18,000 |
| Share of the shared AI gateway | Fixed | $12,000 |
| Service team, 2.5 people | People | $375,000 |
| Policy analyst time for evaluation, 0.3 of a person | People | $30,000 |
| Total run cost | About $501,000 |
Tokens are about 5% of the run cost. People are over 80%. That changes where an architect should spend optimisation effort. Shaving 20% off token spend saves about $5,500 a year. Making the evaluation pipeline efficient enough to save a fifth of the analyst's time saves about $6,000, and designing the system so it needs half a person less to operate saves $75,000.
The table also exposes an earlier decision honestly. The self-hosted guard GPU costs $24,000 a year, almost as much as all the tokens. Section 5 recorded that it was chosen for data protection and capability reasons, not cost. The unit economics confirm that was the right way to describe it.
One more thing to define carefully: what is in a unit cost. Requirement R-NFR-03 in section 3 set "at most 5 cents per policy answer in model and retrieval costs". Meridian's variable cost per question is 2.4 cents, comfortably inside it. Fully loaded, with fixed and people costs spread across all units, the figure is many times higher. Both numbers are true. Write down which one each requirement means, or two teams will argue using different numbers.
Cost controls that matter
Some controls save money; some only look like they do. Meridian ranks them by money saved at its volume.
- Cap output length. Output tokens cost five times input at these prices. A 600-token cap on answers and instructions to be concise matter more than trimming input.
- Prompt caching. The 1,200-token instruction block is identical on every policy answer. Providers that cache repeated prefixes typically charge a small fraction of the normal input price for cached tokens, often around a tenth, though terms vary. At Meridian that saves about 0.3 cents per question, roughly $2,400 a year. Worth doing, because it is free and also lowers latency, but not a strategy.
- Right-sized context. The chunk-count measurement in section 4 cut cost per answer by more than half while improving quality. The best cost savings come with quality gains, not against them.
- Budgets as guardrails. Per-request token caps, a ten-turn limit per conversation, per-user daily limits from the security design, and a monthly budget per capability that alerts at 80%. These protect against runaway cost from bugs and abuse; they are not there to save money in normal use.
Cost against value
A cost figure means little alone. Finance compares it with what it buys. Meridian's pilot measured time saved per unit of work.
| Unit | Per day | Time saved each | Hours saved per day |
|---|---|---|---|
| Policy questions | 3,000 | 2.5 minutes | 125 |
| Summaries for hardship cases | 70 | 9 minutes | 10.5 |
| Summaries for other arrears cases | 530 | 4 minutes | 35 |
| Letters | 70 | 12 minutes | 14 |
| Total | 184.5 |
Over 250 working days, that is about 46,000 staff hours a year. Divide the $501,000 run cost by those hours and the assistant costs about $11 per staff hour saved, against a loaded staff cost of about $40 an hour. That ratio, not cents per call, is the headline number of part A of MER-11.
These unit costs are inputs, not a conclusion. The next lesson turns them into options, a sensitivity table and a case a finance partner can sign.
Check your understanding
0 of 3 answered
1.The pilot estimate was "two cents a call" and the invoice was three times higher. What did the estimate most likely miss?
2.Tokens are about 5% of Meridian's run cost. Where does optimisation effort save the most?
3.A model at a third of the price scored 4 points lower on summary coverage. Why did Meridian keep the original model?