Applied AI Engineering: From Prompt to Production

Course Content

Applied AI Engineering: From Prompt to Production

9 sections · 29 lessons

Choosing a model: the landscape and the trade-offs


The PolicyPal team spent the first two weeks of the project arguing about which model to use. One engineer wanted the largest model because "quality matters for HR". Another wanted an open-weight model on company servers because "HR data must not leave". A third had read a leaderboard. All three had reasons. None of them had evidence about this task.

Public benchmarks measure general skills on public data. PolicyPal's job is narrow: read two or three policy sections and answer one question correctly, briefly and with a citation, for an employee in a specific country. The only benchmark that predicts how a model will do at that job is a small test on that job.

This lesson gives you the landscape, the criteria that matter beyond raw quality, a pilot you can run in an afternoon, and the small client class that keeps the decision reversible.

Thirty real questions, two finalistsLargest hosted model• Passed 28 of 30• About 0.021 dollars per question• 1.4 s to the first token• Won the public leaderboardMid-size model, chosen• Passed 27 of 30• About 0.009 dollars per question• 0.8 s to the first token• Won on this task's numbers
One extra correct answer in thirty did not justify twice the cost and a slower start for every employee.

The landscape in four tiers

Model names change every few months, but the tiers are stable. The prices below are list prices for one provider at the time of writing, in US dollars per million tokens. Other providers sit in similar ranges. Check the current pricing page before you plan a budget.

TierExampleInput / output per 1M tokensTypical use
Large hostedClaude Opus 5$5 / $25Hard reasoning, long agent tasks
Mid hostedClaude Sonnet 5$2 / $10Most product features
Small hostedClaude Haiku 4.5$1 / $5High volume, simple tasks, speed
Open-weight, self-hostedQwen, Llama, Mistral families, 0.5B to 70BYour GPU costData control, narrow tasks, fine-tuning

Two patterns hold across providers. Output tokens cost four to five times as much as input tokens, so a verbose model is an expensive model. And the gap between tiers is smaller on narrow, well-specified tasks than on open-ended ones. A small model that is clearly worse at writing a strategy memo can be nearly as good at "answer from this paragraph".

Open-weight models deserve an honest line. Running one yourself gives you full control of data and no per-token bill, but you now own GPUs, scaling, upgrades and on-call. For a system with 1,200 questions a day, a hosted API is almost always cheaper in total. Self-hosting starts to make sense for narrow, high-volume components, which is exactly where PolicyPal will use it in Section 5.

The criteria beyond quality

Quality on your task is the first filter, but a model that passes it can still be the wrong choice. Check these before you commit.

  • Data handling. Where is data processed, how long does the provider keep it, and is it used for training? HR questions contain health and family details. Harbourline's security team required no training on its data and a documented retention period.
  • Structured output and tool use. PolicyPal needs JSON it can parse and reliable tool calls. Test these directly; they vary more between models than prose quality does.
  • Context window. PolicyPal sends about 3,000 tokens per question, so any current model fits. A document-heavy product might need far more.
  • Latency. Measure time to first token and output speed from your own region, at your own prompt size.
  • Rate limits. Check tokens per minute and requests per minute on your account tier against your peak traffic.
  • Lifecycle. Models are deprecated. Know the provider's notice period, and pin an exact model identifier in configuration.

A pilot that fits in an afternoon

You do not need a big evaluation to shortlist. You need a small, honest one.

  1. Collect 30 real questions — take them from last month's helpdesk tickets, not from your imagination, and include five that the policies do not answer.
  2. Write the gold answer for each — have someone from HR or IT confirm the answer and the policy section it comes from.
  3. Run each candidate with the same prompt and context — same system prompt, same retrieved policy text, three runs per question.
  4. Grade pass or fail — correct, cites the right section, and says "not in the policy" when it should.
  5. Record cost and latency — from each response's token usage and your own timer, not from a price sheet.

Here is what Harbourline's pilot produced, with the relevant policy sections already in the prompt.

CandidatePassed (of 30)Cost per questionMedian time to first token
Large hosted28$0.0211.4 s
Mid hosted27$0.0090.8 s
Small hosted23$0.0040.5 s

The large model won by one question at more than twice the cost and a slower start. The small model was fast and cheap, but four of its seven failures were the same kind of mistake: it answered a UK employee's question with the India policy when both were in the context. The team chose the mid-size model for answers. That is a decision with a reason, a number and a known weakness, and it can be revisited when the eval suite in Section 6 exists.

Keep the model behind a small client

Whatever you choose, you will change it. Prices drop, new models arrive, a provider has an outage. If model calls are spread through the codebase, every change is a refactor. If they go through one small class, a change is a configuration value. This is PolicyPal's client, using the Anthropic Python SDK for the real adapter.

Python
# policypal/llm.pyimport osimport timefrom dataclasses import dataclass, fieldimport anthropic@dataclassclass Reply:    text: str    tool_calls: list[dict] = field(default_factory=list)    stop_reason: str = ""    input_tokens: int = 0    output_tokens: int = 0    latency_ms: int = 0    content: list = field(default_factory=list, repr=False)  # raw blocks, for the agent loopclass LLM:    """The only place in PolicyPal that knows which provider is used."""    def __init__(self, model: str | None = None, timeout_s: float = 30.0):        self.model = model or os.environ.get("POLICYPAL_MODEL", "claude-sonnet-5")        self._client = anthropic.Anthropic(timeout=timeout_s, max_retries=2)    def complete(self, system: str, messages: list[dict], *, tools: list[dict] | None = None,                 schema: dict | None = None, max_tokens: int = 1024) -> Reply:        kwargs = {"model": self.model, "max_tokens": max_tokens,                  "system": system, "messages": messages}        if tools:            kwargs["tools"] = tools        if schema:            kwargs["output_config"] = {"format": {"type": "json_schema", "schema": schema}}        start = time.perf_counter()        resp = self._client.messages.create(**kwargs)        return Reply(            text="".join(b.text for b in resp.content if b.type == "text"),            tool_calls=[{"id": b.id, "name": b.name, "input": b.input}                        for b in resp.content if b.type == "tool_use"],            stop_reason=resp.stop_reason,            input_tokens=resp.usage.input_tokens,            output_tokens=resp.usage.output_tokens,            latency_ms=int((time.perf_counter() - start) * 1000),            content=resp.content,        )

The rest of PolicyPal only ever sees LLM.complete and the Reply dataclass. The model identifier comes from an environment variable, so staging can run a different model from production. The timeout and retry count are set once, here, instead of being forgotten in one of twenty call sites. The raw content blocks are kept because the agent loop in Section 4 must send them back unchanged.

The same shape makes testing easy. A FakeLLM with the same complete method that returns scripted Reply objects lets you test everything around the model without a network call or a bill. You will use one in almost every test file in this course.

Check your understanding

0 of 3 answered

1.The small model was cheapest and fastest but failed 7 of 30 questions, four of them by using the wrong country's policy. What is the best next step if cost pressure is high?

2.Why does PolicyPal read the model identifier from configuration instead of writing it in the code?

3.For Harbourline, when does a self-hosted open-weight model make the most sense?