Course Content
Applied AI Engineering: From Prompt to Production
9 sections · 29 lessons
Choosing a model: the landscape and the trade-offs
The PolicyPal team spent the first two weeks of the project arguing about which model to use. One engineer wanted the largest model because "quality matters for HR". Another wanted an open-weight model on company servers because "HR data must not leave". A third had read a leaderboard. All three had reasons. None of them had evidence about this task.
Public benchmarks measure general skills on public data. PolicyPal's job is narrow: read two or three policy sections and answer one question correctly, briefly and with a citation, for an employee in a specific country. The only benchmark that predicts how a model will do at that job is a small test on that job.
This lesson gives you the landscape, the criteria that matter beyond raw quality, a pilot you can run in an afternoon, and the small client class that keeps the decision reversible.
The landscape in four tiers
Model names change every few months, but the tiers are stable. The prices below are list prices for one provider at the time of writing, in US dollars per million tokens. Other providers sit in similar ranges. Check the current pricing page before you plan a budget.
| Tier | Example | Input / output per 1M tokens | Typical use |
|---|---|---|---|
| Large hosted | Claude Opus 5 | $5 / $25 | Hard reasoning, long agent tasks |
| Mid hosted | Claude Sonnet 5 | $2 / $10 | Most product features |
| Small hosted | Claude Haiku 4.5 | $1 / $5 | High volume, simple tasks, speed |
| Open-weight, self-hosted | Qwen, Llama, Mistral families, 0.5B to 70B | Your GPU cost | Data control, narrow tasks, fine-tuning |
Two patterns hold across providers. Output tokens cost four to five times as much as input tokens, so a verbose model is an expensive model. And the gap between tiers is smaller on narrow, well-specified tasks than on open-ended ones. A small model that is clearly worse at writing a strategy memo can be nearly as good at "answer from this paragraph".
Open-weight models deserve an honest line. Running one yourself gives you full control of data and no per-token bill, but you now own GPUs, scaling, upgrades and on-call. For a system with 1,200 questions a day, a hosted API is almost always cheaper in total. Self-hosting starts to make sense for narrow, high-volume components, which is exactly where PolicyPal will use it in Section 5.
The criteria beyond quality
Quality on your task is the first filter, but a model that passes it can still be the wrong choice. Check these before you commit.
- Data handling. Where is data processed, how long does the provider keep it, and is it used for training? HR questions contain health and family details. Harbourline's security team required no training on its data and a documented retention period.
- Structured output and tool use. PolicyPal needs JSON it can parse and reliable tool calls. Test these directly; they vary more between models than prose quality does.
- Context window. PolicyPal sends about 3,000 tokens per question, so any current model fits. A document-heavy product might need far more.
- Latency. Measure time to first token and output speed from your own region, at your own prompt size.
- Rate limits. Check tokens per minute and requests per minute on your account tier against your peak traffic.
- Lifecycle. Models are deprecated. Know the provider's notice period, and pin an exact model identifier in configuration.
A pilot that fits in an afternoon
You do not need a big evaluation to shortlist. You need a small, honest one.
- Collect 30 real questions — take them from last month's helpdesk tickets, not from your imagination, and include five that the policies do not answer.
- Write the gold answer for each — have someone from HR or IT confirm the answer and the policy section it comes from.
- Run each candidate with the same prompt and context — same system prompt, same retrieved policy text, three runs per question.
- Grade pass or fail — correct, cites the right section, and says "not in the policy" when it should.
- Record cost and latency — from each response's token usage and your own timer, not from a price sheet.
Here is what Harbourline's pilot produced, with the relevant policy sections already in the prompt.
| Candidate | Passed (of 30) | Cost per question | Median time to first token |
|---|---|---|---|
| Large hosted | 28 | $0.021 | 1.4 s |
| Mid hosted | 27 | $0.009 | 0.8 s |
| Small hosted | 23 | $0.004 | 0.5 s |
The large model won by one question at more than twice the cost and a slower start. The small model was fast and cheap, but four of its seven failures were the same kind of mistake: it answered a UK employee's question with the India policy when both were in the context. The team chose the mid-size model for answers. That is a decision with a reason, a number and a known weakness, and it can be revisited when the eval suite in Section 6 exists.
Keep the model behind a small client
Whatever you choose, you will change it. Prices drop, new models arrive, a provider has an outage. If model calls are spread through the codebase, every change is a refactor. If they go through one small class, a change is a configuration value. This is PolicyPal's client, using the Anthropic Python SDK for the real adapter.
1# policypal/llm.py2import os3import time4from dataclasses import dataclass, field56import anthropic78@dataclass9class Reply:10 text: str11 tool_calls: list[dict] = field(default_factory=list)12 stop_reason: str = ""13 input_tokens: int = 014 output_tokens: int = 015 latency_ms: int = 016 content: list = field(default_factory=list, repr=False) # raw blocks, for the agent loop1718class LLM:19 """The only place in PolicyPal that knows which provider is used."""2021 def __init__(self, model: str | None = None, timeout_s: float = 30.0):22 self.model = model or os.environ.get("POLICYPAL_MODEL", "claude-sonnet-5")23 self._client = anthropic.Anthropic(timeout=timeout_s, max_retries=2)2425 def complete(self, system: str, messages: list[dict], *, tools: list[dict] | None = None,26 schema: dict | None = None, max_tokens: int = 1024) -> Reply:27 kwargs = {"model": self.model, "max_tokens": max_tokens,28 "system": system, "messages": messages}29 if tools:30 kwargs["tools"] = tools31 if schema:32 kwargs["output_config"] = {"format": {"type": "json_schema", "schema": schema}}33 start = time.perf_counter()34 resp = self._client.messages.create(**kwargs)35 return Reply(36 text="".join(b.text for b in resp.content if b.type == "text"),37 tool_calls=[{"id": b.id, "name": b.name, "input": b.input}38 for b in resp.content if b.type == "tool_use"],39 stop_reason=resp.stop_reason,40 input_tokens=resp.usage.input_tokens,41 output_tokens=resp.usage.output_tokens,42 latency_ms=int((time.perf_counter() - start) * 1000),43 content=resp.content,44 )The rest of PolicyPal only ever sees LLM.complete and the Reply dataclass. The model identifier comes from an environment variable, so staging can run a different model from production. The timeout and retry count are set once, here, instead of being forgotten in one of twenty call sites. The raw content blocks are kept because the agent loop in Section 4 must send them back unchanged.
The same shape makes testing easy. A FakeLLM with the same complete method that returns scripted Reply objects lets you test everything around the model without a network call or a bill. You will use one in almost every test file in this course.
Check your understanding
0 of 3 answered
1.The small model was cheapest and fastest but failed 7 of 30 questions, four of them by using the wrong country's policy. What is the best next step if cost pressure is high?
2.Why does PolicyPal read the model identifier from configuration instead of writing it in the code?
3.For Harbourline, when does a self-hosted open-weight model make the most sense?