Course Content
Building with LLMs
4 sections · 10 lessons
Adding Structured I/O and Context
An invoice-processing pipeline runs overnight on 3,000 documents. In the morning, 2,847 succeeded and 153 failed. The log for every failure is the same line:
json.decoder.JSONDecodeError: Expecting value: line 1 column 1 (char 0)You pull the raw responses that caused them. Here are four:
1. Sure! Here's the extracted data: {"invoice_number": "INV-9912", "total": 4820.00}2. ```json {"invoice_number": "INV-1043", "total": 917.50} ```3. {"invoice_number": "INV-2288", "total": "1,240.00 GBP"}4. {"invoiceNumber": "INV-7761", "amount_due": 300}Only the first two are parsing failures in the ordinary sense — a friendly preamble and a markdown fence. The third parses fine and then breaks your database, because total is a string with a comma and a currency code where your column expects a number. The fourth parses, has none of the keys you asked for, and writes a row of nulls without raising anything at all.
That is the real shape of the problem. Getting JSON out of a language model is not one problem, it is three: making the response parse, making the fields match, and making the values have the right types. Prompting alone solves the first two most of the time and the third almost never. This lesson is about closing all three properly, and about the other half of the same job — deciding what goes into the context window in the first place when the input is bigger than the window.
Three mechanisms for structured output
1. JSON mode — the raw guarantee
The lowest-level tool. You tell the provider "the output must be syntactically valid JSON", and it constrains generation so it is.
1resp = client.chat.completions.create(2 model=MODEL, # from config, e.g. "gpt-6-luna"3 response_format={"type": "json_object"},4 messages=[5 {"role": "system", "content": "Extract invoice fields. Respond with JSON only."},6 {"role": "user", "content": invoice_text},7 ],8)9data = json.loads(resp.choices[0].message.content) # will not throwThis fixes failures 1 and 2 completely — no preamble, no fence, always parseable. It does nothing about 3 and 4. JSON mode guarantees syntax, not schema. The model can still return whatever keys and types it likes, and it frequently does when the document is unusual.
There is also a sharp edge: with OpenAI's JSON mode, if the word "JSON" does not appear in your prompt, the request errors. And if the model is constrained to emit JSON but has nothing sensible to say, it can generate whitespace until it hits max_tokens — an expensive way to get nothing.
2. Parser-based — instructions in the prompt
Here you define the shape as a Pydantic model, generate format instructions from it, put those in the prompt, and parse and validate the response yourself.
1from pydantic import BaseModel, Field2from langchain_core.output_parsers import PydanticOutputParser3from langchain_core.prompts import ChatPromptTemplate45class Invoice(BaseModel):6 invoice_number: str = Field(description="The invoice reference, e.g. INV-9912")7 total: float = Field(description="Total amount due, as a number with no currency symbol")8 currency: str = Field(description="Three-letter ISO code, e.g. GBP")9 due_date: str = Field(description="Due date in YYYY-MM-DD format")1011parser = PydanticOutputParser(pydantic_object=Invoice)1213prompt = ChatPromptTemplate.from_template(14 "Extract the invoice fields.\n\n{format_instructions}\n\nInvoice:\n{text}"15).partial(format_instructions=parser.get_format_instructions())1617chain = prompt | model | parser18invoice = chain.invoke({"text": invoice_text}) # an Invoice objectThis catches failures 3 and 4 — Pydantic coerces where it can and raises where it cannot, so a wrong key name or an uncoercible value becomes a loud error rather than a silent null. The weakness is that nothing constrains the model; the schema is only advice inside a prompt, so the model can ignore it and you find out at validation time. It works with any provider, including ones with no structured-output feature, which is its main virtue.
3. with_structured_output() — the default you should reach for
This uses the provider's own constrained-decoding or tool-calling machinery under the hood, and hands you a validated object.
1structured = model.with_structured_output(Invoice)2invoice = structured.invoke(f"Extract the invoice fields:\n\n{invoice_text}")34invoice.total # 4820.0 — a float, guaranteed5invoice.currency # "GBP"One line. No format instructions to write, no fence-stripping, no manual parsing. The schema is enforced by the provider rather than requested politely, so all four failures from the opening are structurally impossible. All three major providers now have native schema-constrained output (OpenAI's structured outputs, Anthropic's output_config.format, Gemini's response schema), and with_structured_output uses whichever the model supports.
| JSON mode | Parser-based | with_structured_output() | |
|---|---|---|---|
| Guarantees valid JSON | Yes | No | Yes |
| Guarantees your field names | No | Validates after the fact | Yes |
| Guarantees types | No | Validates and coerces | Yes |
| Prompt tokens spent on schema | Few | Many (full instructions) | Few |
| Works with any provider | Most | All | Most major ones |
| Returns a typed object | No — a dict | Yes | Yes |
| Use when | You need a dict and control the prompt tightly | The provider has no structured-output support | Default |
Valid JSON is not valid data. Syntax guarantees stop your parser crashing; schema guarantees stop your database filling with nulls. They are different problems and only one of them is solved by "respond with JSON".
Designing schemas the model can actually fill in
A schema is a prompt. The field names, the types and especially the descriptions are the clearest instructions you will ever give a model, and small changes to them change extraction quality far more than rewriting the surrounding prose.
1from typing import Literal, Optional2from datetime import date3from pydantic import BaseModel, Field, field_validator45class LineItem(BaseModel):6 description: str = Field(description="What was supplied")7 quantity: float = Field(gt=0, description="Units supplied; 1 if not stated")8 unit_price: float = Field(ge=0, description="Price per unit, excluding tax")9 line_total: float = Field(ge=0, description="quantity x unit_price, excluding tax")1011class Invoice(BaseModel):12 invoice_number: str = Field(description="Reference as printed, e.g. INV-9912")13 issue_date: date = Field(description="Date issued, YYYY-MM-DD")14 due_date: Optional[date] = Field(15 default=None,16 description="Payment due date, YYYY-MM-DD. Null if the invoice does not state one.")17 currency: str = Field(min_length=3, max_length=3,18 description="ISO 4217 code, uppercase, e.g. GBP")19 status: Literal["paid", "unpaid", "overdue", "unknown"] = Field(20 description="Payment status. Use 'unknown' if the invoice does not say.")21 line_items: list[LineItem] = Field(default_factory=list)22 subtotal: float = Field(ge=0, description="Sum of line totals, excluding tax")23 tax: float = Field(ge=0, description="Total tax charged")24 total: float = Field(ge=0, description="Amount due including tax")2526 @field_validator("currency")27 @classmethod28 def upper(cls, v: str) -> str:29 return v.upper()Five design decisions in there, each fixing a specific real failure:
| Decision | Failure it prevents |
|---|---|
Literal[...] for status | The model inventing "pending", "part-paid", "PAID" |
Optional with an explicit "null if absent" description | The model inventing a plausible due date that is not on the document |
Constraints (gt=0, ge=0, length 3) | Negative quantities and "British Pounds" in the currency field |
| Descriptions naming the format | "1st March 2026", "03/01/26" and "2026-03-01" all appearing in one column |
Nested LineItem list | Line items flattened into one string you then have to re-parse |
The optional-field lesson deserves emphasis, because it is the most damaging silent failure in extraction work. A model handed a required field it cannot find will produce something anyway — that is what generation does. Making the field optional and saying explicitly "null if not stated" gives it a legitimate way to say "not present", and the difference in hallucinated values is dramatic.
Two limits to respect. Deeply nested schemas — four or five levels — degrade extraction quality noticeably; flatten where you can. And the schema itself is sent as tokens on every request, so a 60-field model with long descriptions is a real per-call cost. Extract what you will use.
Counting tokens, because the window is finite
Structured output handles what comes out. The other half of the job is what goes in — and inputs regularly exceed what the model will accept.
You cannot manage a budget you cannot measure. For OpenAI models, tiktoken counts locally, offline and free:
1import tiktoken23# Newer model names are not always in tiktoken's lookup table, so name the4# encoding directly. o200k_base is the newest public one: a close estimate.5enc = tiktoken.get_encoding("o200k_base")67def count(text: str) -> int:8 return len(enc.encode(text))910count("The quick brown fox jumps over the lazy dog.") # 1011count(invoice_text) # 1,842Three things to know about this. First, tokenisers are provider-specific — tiktoken gives OpenAI's counts and is only an approximation elsewhere. OpenAI (client.responses.input_tokens.count) and Anthropic (client.messages.count_tokens) both expose exact counting endpoints, and Google exposes client.models.count_tokens; use the provider's own counter for the provider you are billing against, or your budget arithmetic will be quietly wrong by 10–35%.
Second, message overhead is real. Each message in a chat request carries a few tokens of role and delimiter framing — roughly 3 to 4 per message, plus a few for the reply priming. Counting only the content of a 30-message conversation undercounts by around 100 tokens. Small, but it is exactly the margin that makes a request that "should fit" fail.
Third, the useful rules of thumb: about 4 characters per token for English prose, about 0.75 words per token. So 1,000 words is roughly 1,330 tokens, and a 200-page document at 500 words a page is roughly 133,000 tokens. Use these for planning; use a real counter for anything you enforce.
1def fits(system: str, history: list, question: str, model_window: int,2 reserve_for_reply: int = 1500) -> bool:3 used = count(system) + count(question) + 44 used += sum(count(m.content) + 4 for m in history)5 return used + reserve_for_reply <= model_windowThe reserve_for_reply parameter is the piece people forget. The context window holds input and output. Filling 127,500 tokens of a 128,000-token window leaves 500 tokens for the answer, and you get a response cut off mid-sentence with no error — the same silent truncation that produces "the model got worse" bug reports.
Always reserve output space before you fill the context window. A request that fits and an answer that fits are two different constraints, and only one of them fails loudly.
Fitting more in than fits: three strategies
Sliding window — keep the most recent
1def window(messages, budget_tokens: int):2 kept, used = [], 03 for m in reversed(messages): # newest first4 c = count(m.content) + 45 if used + c > budget_tokens:6 break7 kept.append(m); used += c8 kept.reverse()9 while kept and kept[0].type == "ai": # never start on an assistant turn10 kept.pop(0)11 return keptCheap, deterministic, no extra API call. The cost is total amnesia about anything that falls out — the customer's name mentioned in turn two vanishes at turn thirty. Right for chat where recency dominates, wrong when early facts matter.
Summarisation — compress the old
1SUMMARISE = ChatPromptTemplate.from_template(2 "Compress the exchange below into at most 150 words. Preserve every name, "3 "identifier, number, date and commitment exactly. Drop pleasantries and "4 "restatements. Write in the third person.\n\n{chunk}")56def compress(messages, keep_recent=8):7 if len(messages) <= keep_recent + 4:8 return messages9 old, recent = messages[:-keep_recent], messages[-keep_recent:]10 text = "\n".join(f"{m.type}: {m.content}" for m in old)11 summary = (SUMMARISE | cheap_model | StrOutputParser()).invoke({"chunk": text})12 return [SystemMessage(content=f"Summary of earlier conversation:\n{summary}")] + recentThe arithmetic that justifies it: 24 old messages averaging 120 tokens is 2,880 tokens, compressed to about 200 — a 93% reduction on that portion. The costs are an extra API call, a few hundred milliseconds, and irreversible information loss. The explicit instruction to preserve identifiers is what stops the loss being catastrophic; a summariser told merely to "summarise" reliably drops the order number, which is precisely the token you needed.
Retrieval — fetch only what is relevant
The strongest option when the source is a corpus rather than a conversation. Instead of choosing which parts of a 300-page document to keep, index it in chunks and retrieve the four chunks that match the question. Context stays roughly constant no matter how large the corpus grows.
| Strategy | Extra cost | Loses information | Best for |
|---|---|---|---|
| Sliding window | None | Yes, silently | Chat where only recency matters |
| Summarisation | One cheap call per compression | Yes, controllably | Long conversations with facts worth keeping |
| Retrieval | Embedding + index infrastructure | Only what is not retrieved | Large static corpora, document Q&A |
| Structured state alongside | Negligible | No, for tracked fields | Facts your code depends on being exact |
These combine, and in production they usually should: a structured state object for the handful of facts that must never be lost, a summary of the middle, the last few turns verbatim, and retrieval for the reference material.
A complete extraction, end to end
Pulling both halves together — bounded input, guaranteed output. The task: extract action items from a meeting transcript that may be far too long for one request.
1from typing import Literal, Optional2from pydantic import BaseModel, Field34class ActionItem(BaseModel):5 task: str = Field(description="What must be done, as an imperative phrase")6 owner: Optional[str] = Field(7 default=None, description="Person responsible. Null if not assigned aloud.")8 due: Optional[str] = Field(9 default=None, description="Deadline as YYYY-MM-DD. Null if none was stated.")10 priority: Literal["high", "medium", "low"] = Field(11 description="Inferred urgency from the discussion")12 quote: str = Field(description="The sentence from the transcript that states this item")1314class Extraction(BaseModel):15 items: list[ActionItem] = Field(default_factory=list)16 decisions: list[str] = Field(default_factory=list,17 description="Decisions reached, one sentence each")1819extractor = model.with_structured_output(Extraction)2021def extract(transcript: str, window_tokens: int = 100_000) -> Extraction:22 if count(transcript) <= window_tokens:23 return extractor.invoke(PROMPT.format(text=transcript))2425 # Too long: process in overlapping chunks and merge26 merged = Extraction()27 for chunk in chunk_by_tokens(transcript, size=window_tokens, overlap=2_000):28 part = extractor.invoke(PROMPT.format(text=chunk))29 merged.items.extend(part.items)30 merged.decisions.extend(part.decisions)3132 seen, unique = set(), []33 for it in merged.items:34 key = it.task.lower().strip()35 if key not in seen:36 seen.add(key); unique.append(it)37 merged.items = unique38 return mergedThree deliberate choices. The quote field forces the model to ground each item in actual transcript text, which makes hallucinated action items both rarer and instantly checkable — you can assert that item.quote appears in the source and reject the ones that do not. The 2,000-token overlap between chunks stops an action item that straddles a boundary from being lost by both sides. And the deduplication step exists because of the overlap: items in the shared region get extracted twice, and merging without dedup silently doubles them.
Where people get this wrong
Believing JSON mode enforces a schema. It enforces syntax. Failures 3 and 4 from the opening pass JSON mode cleanly and corrupt your data anyway.
Making every field required. The fastest way to manufacture hallucinations. If the document might not contain a due date, the schema must permit its absence — otherwise the model fabricates one, and it will be plausible enough that nobody notices for months.
Writing schemas without descriptions. A bare due_date: str gets you three date formats in one column. The description is the instruction; the field name alone is not.
Estimating tokens with len(text) / 4 and then enforcing a limit with it. Fine for a rough plan, wrong for a gate. Code, non-English text and long identifiers all tokenise far worse than the rule of thumb — code is often closer to 2.5 characters per token — so the estimate under-counts exactly where the input is largest.
Forgetting the reply budget. Filling the window to the brim gives you a truncated answer and no exception.
Using one provider's tokeniser to budget another provider's calls. Different vocabularies, different counts, wrong bill.
Chunking without overlap. Every fact that crosses a chunk boundary is lost, and it is lost silently — the extraction looks complete.
What this changes about how you build
Structured output moves the failure from runtime to design time, and that is the whole point. Before it, a model returning "PAID" instead of "paid" is a production incident discovered by a downstream report; after it, that value cannot be produced at all, because Literal["paid", "unpaid", "overdue", "unknown"] forbids it. You have converted a class of bug into a class of impossibility.
The practical consequence is that the schema becomes the specification of the feature. When someone asks for a new field, the conversation is about the type, whether it is optional, what its allowed values are and how to describe it — which is a far more productive conversation than "make the prompt better". Write the schema first, before the prompt, and the prompt often turns out to be two sentences.
The context side has a matching consequence: measure before you optimise. Log the input token count of every call from day one. Almost every team that hits a context limit discovers the cause is not the user's document but their own accumulated system prompt, a tool returning raw JSON, or a history that was never trimmed. Those are cheap fixes, but only if you have the number that points at them. Without it you will reach for a bigger model, pay several times more, and hit the same wall three months later with the same underlying cause.