Course Content
Building with LLMs
4 sections · 10 lessons
OpenAI vs. Anthropic vs. Gemini — Choosing a Model Provider
A team builds a contract-review feature. They pick a provider on day one — whichever one they already had an account with — and wire it in directly: the SDK client, the message format, the tool-calling shape, the way errors come back. Six weeks later the feature works.
Then two things happen in the same fortnight. First, a customer sends a 400-page master services agreement, and the model rejects it because the whole document does not fit in the context window. Second, the provider has a four-hour outage and the feature returns 503 to every user, because there is nothing else it can call.
Both problems were decidable on day one, and both were decided by accident. The team never asked what the model needed to be good at, and never asked what happens when it is unavailable.
This lesson is about making that choice deliberately. Three providers dominate: OpenAI, Anthropic and Google. They are more alike than the marketing suggests — all three take a list of messages, return text, support tool calling and structured output, and charge per token. The differences that matter are narrower and more practical than "which is smarter", and they are what you can actually design around.
The shape all three share
Before the differences, get the sameness clear, because it is what makes switching feasible at all.
Every one of these APIs is a function from a conversation to a message. You send a system instruction plus an alternating list of user and assistant turns; you get back an assistant turn. You are billed on input tokens (everything you sent, re-sent in full on every request) and output tokens (what the model generated). A token is roughly three-quarters of an English word — 1,000 tokens is about 750 words.
The important consequence of "re-sent in full on every request" is that a conversation gets more expensive with every turn, because turn ten sends turns one through nine again. That is true on all three providers and it surprises people every time.
Every provider also ships roughly three tiers, and the tier matters more than the brand:
| Tier | What it is for | Rough relative cost | Typical use |
|---|---|---|---|
| Small / fast | High volume, simple judgement | 1× | Classification, routing, tagging, extraction from clean text |
| Mid | The general workhorse | 2–20× | Chat, summarisation, most RAG answering, tool use |
| Large / reasoning | Hard multi-step problems | 5–100× | Code generation, long agentic runs, contract analysis, planning |
The spread between tiers differs by provider. On Anthropic's published prices the large tier costs five times the small one; on OpenAI's GPT-6 prices it is a hundred times. Either way, the tier is the first number to get right.
Choosing the right tier within one provider changes your bill by 10× or more. Choosing between providers at the same tier usually changes it by less than 2×. Get the tier right before you argue about the brand.
OpenAI
The most widely deployed, with the largest ecosystem of tutorials, wrappers and third-party integrations. If a tool claims to support "an LLM", it supports OpenAI first.
1import os2from openai import OpenAI34client = OpenAI() # reads OPENAI_API_KEY from the environment5MODEL = os.environ.get("OPENAI_MODEL", "gpt-6-luna")67resp = client.responses.create(8 model=MODEL,9 instructions="You are a terse contract analyst.",10 input="What is the notice period in this clause? ...",11)12print(resp.output_text)This is the Responses API, which OpenAI now points new work at. The older Chat Completions endpoint (client.chat.completions.create) still works and fills most existing code and tutorials, but on the GPT-6 models tool calling with reasoning switched on needs Responses.
Tool calling
Tools are declared as JSON Schema. The model does not run them; it returns a function_call item, you execute it, and you send the result back as a function_call_output item.
1tools = [{2 "type": "function",3 "name": "get_contract_expiry",4 "description": "Return the expiry date for a contract by its ID.",5 "parameters": {6 "type": "object",7 "properties": {"contract_id": {"type": "string"}},8 "required": ["contract_id"],9 },10}]1112resp = client.responses.create(model=MODEL, input=messages, tools=tools)1314call = next(item for item in resp.output if item.type == "function_call")15result = get_contract_expiry(**json.loads(call.arguments))1617messages += resp.output # keep the model's request in the history18messages.append({"type": "function_call_output",19 "call_id": call.call_id, "output": result})Structured output
OpenAI can constrain generation to a schema, which means the response is guaranteed to parse rather than merely likely to:
1from pydantic import BaseModel23class Clause(BaseModel):4 clause_type: str5 notice_days: int6 is_auto_renewing: bool78resp = client.responses.parse(9 model=MODEL,10 input=messages,11 text_format=Clause,12)13clause = resp.output_parsed # a real Clause objectReach for OpenAI when you need the broadest third-party compatibility, when you want image generation and speech in the same account, or when the team already knows the SDK and the task is unremarkable.
Anthropic (Claude)
Claude's distinguishing traits in practice are a very large context window across the whole current line, strong instruction-following on long structured documents, and thinking that the model manages itself rather than one you have to budget.
1from anthropic import Anthropic23client = Anthropic() # reads ANTHROPIC_API_KEY45resp = client.messages.create(6 model="claude-opus-5",7 max_tokens=4096,8 system="You are a terse contract analyst.",9 messages=[{"role": "user", "content": "What is the notice period in this clause? ..."}],10)11print(resp.content[0].text)Three shape differences catch people migrating from OpenAI. The system prompt is a top-level parameter, not a message with role="system". max_tokens is required, not optional — leave it out and the request fails. Setting it too low is the more common bug: the response is truncated mid-sentence, the only signal is stop_reason == "max_tokens", and you spend an hour blaming the prompt. And Claude Opus 5 and Sonnet 5 reject temperature and top_p, so do not carry those over from old code.
Structured output works the same way as OpenAI's: client.messages.parse(model=..., max_tokens=..., messages=..., output_format=Clause) returns a validated object in resp.parsed_output.
Long-context document analysis
The current Claude models carry a 1M-token context window (Haiku 4.5 is the exception at 200K). One million tokens is roughly 750,000 words — the entire 400-page agreement, plus every amendment, plus your instructions, in a single request with no chunking.
1contract = open("msa_2026.txt").read() # ~310,000 tokens23resp = client.messages.create(4 model="claude-opus-5",5 max_tokens=8000,6 system="You extract obligations from commercial contracts. Cite clause numbers.",7 messages=[{"role": "user", "content": f"<contract>{contract}</contract>\n\n"8 "List every obligation that falls on the supplier."}],9)Two practical notes. Claude responds well to XML-style delimiters like the <contract> tags above — they make the boundary between data and instruction unambiguous, which matters enormously when the data is 300,000 tokens long. And a request this size is expensive to repeat: at 5 dollars per million input tokens, 310,000 tokens costs about 1.55 dollars per question. If you plan to ask twenty questions about the same contract, use prompt caching so the document is charged at the cached rate after the first call, rather than re-billing 1.55 dollars twenty times.
Tool use
1tools = [{2 "name": "get_contract_expiry",3 "description": "Return the expiry date for a contract by its ID.",4 "input_schema": {5 "type": "object",6 "properties": {"contract_id": {"type": "string"}},7 "required": ["contract_id"],8 },9}]1011resp = client.messages.create(12 model="claude-opus-5", max_tokens=1024, tools=tools, messages=messages)1314if resp.stop_reason == "tool_use":15 block = next(b for b in resp.content if b.type == "tool_use")16 result = get_contract_expiry(**block.input)17 messages += [18 {"role": "assistant", "content": resp.content},19 {"role": "user", "content": [{20 "type": "tool_result", "tool_use_id": block.id, "content": result}]},21 ]The differences from OpenAI are mechanical, not conceptual: the schema key is input_schema rather than parameters, you branch on stop_reason == "tool_use", and tool results go back as a user message containing tool_result blocks rather than as a function_call_output item. Every provider-agnostic wrapper you will ever read is mostly code that translates between these three conventions.
Extended thinking
Claude can reason before answering. On the current models this is adaptive: the model decides how much thinking each request deserves, rather than you guessing a token budget in advance. Opus 5 thinks adaptively by default; the thinking line below just makes that explicit. (Haiku 4.5 is the exception and still takes a fixed budget_tokens.)
1resp = client.messages.create(2 model="claude-opus-5",3 max_tokens=16000,4 thinking={"type": "adaptive"},5 output_config={"effort": "high"},6 messages=[{"role": "user", "content": "Which of these three indemnity clauses "7 "exposes us most, and quantify why."}],8)Thinking tokens are billed as output. That is the trade: a request that would have produced 400 output tokens might now produce 3,000, so the same question costs several times more but is materially more likely to be right on genuinely hard problems. effort is the dial — low for routine subtasks, high (the default) for anything intelligence-sensitive, max when correctness matters more than cost. Running a classification task at high effort is money set on fire; running a multi-clause legal comparison at low effort is accuracy set on fire.
Reasoning is not free intelligence. It converts output tokens into accuracy at a fixed exchange rate, and it is only worth it when a wrong answer costs more than the extra tokens.
Reach for Claude when the input is long and structured, when you need careful instruction-following over documents, or when the task benefits from real reasoning.
Google Gemini
Gemini's distinguishing traits are native multimodality — video and audio go straight into the same request as text, not through a separate transcription step — and deep integration with Google Cloud if you already live there.
1import os2from google import genai34client = genai.Client() # reads GEMINI_API_KEY or GOOGLE_API_KEY5MODEL = os.environ.get("GEMINI_MODEL", "gemini-3.8-flash")67resp = client.models.generate_content(8 model=MODEL,9 contents="What is the notice period in this clause? ...",10 config={"system_instruction": "You are a terse contract analyst."},11)12print(resp.text)Multimodal input
This is the capability with the least equivalent elsewhere. A one-hour recorded meeting can go into the request directly:
1video = client.files.upload(file="board_meeting.mp4")23resp = client.models.generate_content(4 model=MODEL,5 contents=[video, "List every decision made, with the timestamp and who proposed it."],6)Doing the same thing on a text-only model means transcribing first, which loses the slides, the whiteboard, and who was actually in the room. If your data is video, audio or dense scanned imagery, that difference is not a nicety — it is the whole feature.
Tool calling
1def get_contract_expiry(contract_id: str) -> str:2 """Return the expiry date for a contract by its ID."""3 return lookup(contract_id)45resp = client.models.generate_content(6 model=MODEL,7 contents="When does contract C-8891 expire?",8 config={"tools": [get_contract_expiry]},9)Gemini can take plain Python functions and derive the schema from the signature and docstring, which is the least ceremonious of the three. The same rule applies as everywhere: the docstring is the tool description the model reads, so a vague one produces a tool the model never calls.
Reach for Gemini when your inputs are video or audio, when you are already on Google Cloud and want IAM and billing to be one thing, or when you want a cheap fast tier with a large context window for high-volume work.
Costing a real feature, with the arithmetic
Abstract price-per-million comparisons mislead, because your bill is driven by your token shape, not by the headline rate. Work an actual case.
The feature: a support assistant. 50,000 requests a month. Each request sends a 900-token system prompt, about 1,600 tokens of retrieved context, and a 100-token question — 2,600 input tokens — and generates a 350-token answer.
Monthly totals: 50,000 × 2,600 = 130 million input tokens and 50,000 × 350 = 17.5 million output tokens.
Now price it against three published Claude rates (as of September 2026), which give a clean illustration of how tier choice dominates:
| Model | Input per 1M | Output per 1M | Input cost | Output cost | Monthly total |
|---|---|---|---|---|---|
| Haiku 4.5 | 1.00 | 5.00 | 130.00 | 87.50 | 217.50 |
| Sonnet 5 | 2.00 | 10.00 | 260.00 | 175.00 | 435.00 |
| Opus 5 | 5.00 | 25.00 | 650.00 | 437.50 | 1,087.50 |
All figures in US dollars per month. Check the arithmetic on one row: 130 million input tokens is 130 units of a million, so at 2.00 dollars per million that is 130 × 2.00 = 260.00 dollars; 17.5 × 10.00 = 175.00 dollars; total 435.00 dollars. Prices change, so re-check the provider's pricing page before you rely on a table like this.
Two lessons fall straight out. First, the small tier is five times cheaper than the large one for identical traffic — that gap dwarfs any difference between vendors at the same tier, which is why "which provider is cheapest" is usually the wrong question. Second, this workload is input-heavy: 130 million in against 17.5 million out. Trimming 600 tokens off that system prompt saves 50,000 × 600 = 30 million input tokens, which is 60 dollars a month on Sonnet 5 — more than any output-side optimisation you could make. Optimise the side that dominates your ratio, and measure the ratio before you optimise anything.
One more number worth having: prompt caching. If the 900-token system prompt and a stable 1,600-token context are cached, the cached portion is billed at a small fraction of the input rate on repeat calls. On an input-heavy workload like this one, that is the single largest lever available, and it requires no change to the model or the prompt — only to where you put the cache breakpoints.
Choosing, in order
Work through these in sequence and stop at the first one that binds. Most decisions are settled by question two or three.
| Question | If yes |
|---|---|
| Is the input video or audio? | Gemini — the others need a separate transcription stage |
| Does a single request need more than ~200K tokens? | Claude Opus 5 and Sonnet 5 and OpenAI's GPT-6 models take about 1M; check the specific model's window (Claude Haiku 4.5 stops at 200K) |
| Is the task hard multi-step reasoning where being wrong is expensive? | A reasoning-capable large tier, and switch thinking on |
| Is it high-volume simple judgement (classify, route, tag)? | The smallest tier of any provider — differences are marginal |
| Are you standardised on one cloud with strict data-residency rules? | That cloud's offering, or the provider available through it |
| None of the above binds? | Use what your team already knows. The difference will not be your bottleneck. |
Not being trapped: the provider-agnostic layer
The contract-review team's second failure was structural. Provider SDK calls were scattered through the codebase, so switching meant touching every file, and there was no second path when the first one broke.
The fix is one interface with one implementation per provider. Keep it small — this is not a framework, it is about forty lines:
1from abc import ABC, abstractmethod23class LLMProvider(ABC):4 @abstractmethod5 def complete(self, system: str, user: str, max_tokens: int = 2048) -> str: ...67class AnthropicProvider(LLMProvider):8 def __init__(self, model="claude-opus-5"):9 from anthropic import Anthropic10 self.client, self.model = Anthropic(), model1112 def complete(self, system, user, max_tokens=2048):13 r = self.client.messages.create(14 model=self.model, max_tokens=max_tokens,15 system=system, messages=[{"role": "user", "content": user}])16 return r.content[0].text1718class OpenAIProvider(LLMProvider):19 def __init__(self, model="gpt-6-luna"):20 from openai import OpenAI21 self.client, self.model = OpenAI(), model2223 def complete(self, system, user, max_tokens=2048):24 r = self.client.responses.create(25 model=self.model, max_output_tokens=max_tokens,26 instructions=system, input=user)27 return r.output_textThen fall back when the primary fails:
1import logging23class FallbackProvider(LLMProvider):4 def __init__(self, *providers):5 self.providers = providers67 def complete(self, system, user, max_tokens=2048):8 last = None9 for p in self.providers:10 try:11 return p.complete(system, user, max_tokens)12 except Exception as exc:13 logging.warning("provider %s failed: %s", type(p).__name__, exc)14 last = exc15 raise RuntimeError(f"all providers failed; last error: {last}")1617llm = FallbackProvider(AnthropicProvider(), OpenAIProvider())Three honest caveats, because naive fallback code causes its own outages.
Do not fall back on 4xx errors. A 400 means your request was malformed and the second provider will reject it too — you have just doubled your latency and your error rate. Retry only on timeouts, connection errors, 429 and 5xx.
The abstraction leaks, and pretending otherwise is worse than not having it. Tool-calling shapes differ, structured-output mechanisms differ, streaming event formats differ, and one provider's reasoning parameters have no equivalent elsewhere. Keep the shared interface to what genuinely is shared, and let the specialist paths be explicit.
Prompts are not portable. A prompt tuned against one model will usually work on another and occasionally will not — different models react differently to formatting, to role instructions, to few-shot examples. If you fail over silently, you are serving untested prompt-model pairs to users. Run your evaluation set against every provider in the chain before you ship it, not after the incident.
What this means when you build
The decisions that actually bite are these, in this order.
Choose the tier before the vendor. The 5× cost gap between the small and large tier of one provider is larger than the gap between providers at the same tier. Start every feature on the smallest tier that could plausibly work, build an evaluation set of 50 real inputs with known-correct answers, and only move up when you can point at the specific cases the small model gets wrong.
Route per task, not per application. There is no reason one application must use one model. Send intent classification to a small fast model, retrieval-augmented answering to a mid tier, and the rare hard analysis to a large reasoning model. The provider-agnostic interface above is what makes that a configuration change rather than a refactor.
Measure your token ratio on day one. Log input and output tokens for every call from the very first commit. Without that number you cannot tell whether to shorten prompts, shorten answers, add caching, or change tier — and every one of those is a different piece of work.
Assume the outage. Every provider goes down. Decide in advance what your feature does when it happens: fail over to a second provider, serve a degraded cached answer, or fail loudly with a clear message. Any of the three is a legitimate choice. Discovering that you have no choice at 3 a.m. is not.