Building with LLMs

Adding Structured I/O and Context


An invoice-processing pipeline runs overnight on 3,000 documents. In the morning, 2,847 succeeded and 153 failed. The log for every failure is the same line:

Text
json.decoder.JSONDecodeError: Expecting value: line 1 column 1 (char 0)

You pull the raw responses that caused them. Here are four:

Text
1.  Sure! Here's the extracted data:    {"invoice_number": "INV-9912", "total": 4820.00}2.  ```json    {"invoice_number": "INV-1043", "total": 917.50}    ```3.  {"invoice_number": "INV-2288", "total": "1,240.00 GBP"}4.  {"invoiceNumber": "INV-7761", "amount_due": 300}

Only the first two are parsing failures in the ordinary sense — a friendly preamble and a markdown fence. The third parses fine and then breaks your database, because total is a string with a comma and a currency code where your column expects a number. The fourth parses, has none of the keys you asked for, and writes a row of nulls without raising anything at all.

That is the real shape of the problem. Getting JSON out of a language model is not one problem, it is three: making the response parse, making the fields match, and making the values have the right types. Prompting alone solves the first two most of the time and the third almost never. This lesson is about closing all three properly, and about the other half of the same job — deciding what goes into the context window in the first place when the input is bigger than the window.

Three mechanisms for getting a filled schema backyesnoyoursany modelnoyesyoursany modelyesyesbuilt insupportedonlyValid JSONSchema checkedRetry logicPortabilityJSON modePrompt parserwith_structured_outputJSON mode promises syntax, not fields — a well-formed object with the wrong keys still parses.
Only the third row gives you both guarantees, which is why the other two exist mainly as fallbacks.

Three mechanisms for structured output

1. JSON mode — the raw guarantee

The lowest-level tool. You tell the provider "the output must be syntactically valid JSON", and it constrains generation so it is.

Python
resp = client.chat.completions.create(    model=MODEL,                  # from config, e.g. "gpt-6-luna"    response_format={"type": "json_object"},    messages=[        {"role": "system", "content": "Extract invoice fields. Respond with JSON only."},        {"role": "user", "content": invoice_text},    ],)data = json.loads(resp.choices[0].message.content)   # will not throw

This fixes failures 1 and 2 completely — no preamble, no fence, always parseable. It does nothing about 3 and 4. JSON mode guarantees syntax, not schema. The model can still return whatever keys and types it likes, and it frequently does when the document is unusual.

There is also a sharp edge: with OpenAI's JSON mode, if the word "JSON" does not appear in your prompt, the request errors. And if the model is constrained to emit JSON but has nothing sensible to say, it can generate whitespace until it hits max_tokens — an expensive way to get nothing.

2. Parser-based — instructions in the prompt

Here you define the shape as a Pydantic model, generate format instructions from it, put those in the prompt, and parse and validate the response yourself.

Python
from pydantic import BaseModel, Fieldfrom langchain_core.output_parsers import PydanticOutputParserfrom langchain_core.prompts import ChatPromptTemplateclass Invoice(BaseModel):    invoice_number: str = Field(description="The invoice reference, e.g. INV-9912")    total: float = Field(description="Total amount due, as a number with no currency symbol")    currency: str = Field(description="Three-letter ISO code, e.g. GBP")    due_date: str = Field(description="Due date in YYYY-MM-DD format")parser = PydanticOutputParser(pydantic_object=Invoice)prompt = ChatPromptTemplate.from_template(    "Extract the invoice fields.\n\n{format_instructions}\n\nInvoice:\n{text}").partial(format_instructions=parser.get_format_instructions())chain = prompt | model | parserinvoice = chain.invoke({"text": invoice_text})     # an Invoice object

This catches failures 3 and 4 — Pydantic coerces where it can and raises where it cannot, so a wrong key name or an uncoercible value becomes a loud error rather than a silent null. The weakness is that nothing constrains the model; the schema is only advice inside a prompt, so the model can ignore it and you find out at validation time. It works with any provider, including ones with no structured-output feature, which is its main virtue.

3. with_structured_output() — the default you should reach for

This uses the provider's own constrained-decoding or tool-calling machinery under the hood, and hands you a validated object.

Python
structured = model.with_structured_output(Invoice)invoice = structured.invoke(f"Extract the invoice fields:\n\n{invoice_text}")invoice.total      # 4820.0  — a float, guaranteedinvoice.currency   # "GBP"

One line. No format instructions to write, no fence-stripping, no manual parsing. The schema is enforced by the provider rather than requested politely, so all four failures from the opening are structurally impossible. All three major providers now have native schema-constrained output (OpenAI's structured outputs, Anthropic's output_config.format, Gemini's response schema), and with_structured_output uses whichever the model supports.

JSON modeParser-basedwith_structured_output()
Guarantees valid JSONYesNoYes
Guarantees your field namesNoValidates after the factYes
Guarantees typesNoValidates and coercesYes
Prompt tokens spent on schemaFewMany (full instructions)Few
Works with any providerMostAllMost major ones
Returns a typed objectNo — a dictYesYes
Use whenYou need a dict and control the prompt tightlyThe provider has no structured-output supportDefault

Valid JSON is not valid data. Syntax guarantees stop your parser crashing; schema guarantees stop your database filling with nulls. They are different problems and only one of them is solved by "respond with JSON".

Designing schemas the model can actually fill in

A schema is a prompt. The field names, the types and especially the descriptions are the clearest instructions you will ever give a model, and small changes to them change extraction quality far more than rewriting the surrounding prose.

Python
from typing import Literal, Optionalfrom datetime import datefrom pydantic import BaseModel, Field, field_validatorclass LineItem(BaseModel):    description: str = Field(description="What was supplied")    quantity: float = Field(gt=0, description="Units supplied; 1 if not stated")    unit_price: float = Field(ge=0, description="Price per unit, excluding tax")    line_total: float = Field(ge=0, description="quantity x unit_price, excluding tax")class Invoice(BaseModel):    invoice_number: str = Field(description="Reference as printed, e.g. INV-9912")    issue_date: date = Field(description="Date issued, YYYY-MM-DD")    due_date: Optional[date] = Field(        default=None,        description="Payment due date, YYYY-MM-DD. Null if the invoice does not state one.")    currency: str = Field(min_length=3, max_length=3,                          description="ISO 4217 code, uppercase, e.g. GBP")    status: Literal["paid", "unpaid", "overdue", "unknown"] = Field(        description="Payment status. Use 'unknown' if the invoice does not say.")    line_items: list[LineItem] = Field(default_factory=list)    subtotal: float = Field(ge=0, description="Sum of line totals, excluding tax")    tax: float = Field(ge=0, description="Total tax charged")    total: float = Field(ge=0, description="Amount due including tax")    @field_validator("currency")    @classmethod    def upper(cls, v: str) -> str:        return v.upper()

Five design decisions in there, each fixing a specific real failure:

DecisionFailure it prevents
Literal[...] for statusThe model inventing "pending", "part-paid", "PAID"
Optional with an explicit "null if absent" descriptionThe model inventing a plausible due date that is not on the document
Constraints (gt=0, ge=0, length 3)Negative quantities and "British Pounds" in the currency field
Descriptions naming the format"1st March 2026", "03/01/26" and "2026-03-01" all appearing in one column
Nested LineItem listLine items flattened into one string you then have to re-parse

The optional-field lesson deserves emphasis, because it is the most damaging silent failure in extraction work. A model handed a required field it cannot find will produce something anyway — that is what generation does. Making the field optional and saying explicitly "null if not stated" gives it a legitimate way to say "not present", and the difference in hallucinated values is dramatic.

Two limits to respect. Deeply nested schemas — four or five levels — degrade extraction quality noticeably; flatten where you can. And the schema itself is sent as tokens on every request, so a 60-field model with long descriptions is a real per-call cost. Extract what you will use.

Counting tokens, because the window is finite

Structured output handles what comes out. The other half of the job is what goes in — and inputs regularly exceed what the model will accept.

You cannot manage a budget you cannot measure. For OpenAI models, tiktoken counts locally, offline and free:

Python
import tiktoken# Newer model names are not always in tiktoken's lookup table, so name the# encoding directly. o200k_base is the newest public one: a close estimate.enc = tiktoken.get_encoding("o200k_base")def count(text: str) -> int:    return len(enc.encode(text))count("The quick brown fox jumps over the lazy dog.")   # 10count(invoice_text)                                      # 1,842

Three things to know about this. First, tokenisers are provider-specific — tiktoken gives OpenAI's counts and is only an approximation elsewhere. OpenAI (client.responses.input_tokens.count) and Anthropic (client.messages.count_tokens) both expose exact counting endpoints, and Google exposes client.models.count_tokens; use the provider's own counter for the provider you are billing against, or your budget arithmetic will be quietly wrong by 10–35%.

Second, message overhead is real. Each message in a chat request carries a few tokens of role and delimiter framing — roughly 3 to 4 per message, plus a few for the reply priming. Counting only the content of a 30-message conversation undercounts by around 100 tokens. Small, but it is exactly the margin that makes a request that "should fit" fail.

Third, the useful rules of thumb: about 4 characters per token for English prose, about 0.75 words per token. So 1,000 words is roughly 1,330 tokens, and a 200-page document at 500 words a page is roughly 133,000 tokens. Use these for planning; use a real counter for anything you enforce.

Python
def fits(system: str, history: list, question: str, model_window: int,         reserve_for_reply: int = 1500) -> bool:    used = count(system) + count(question) + 4    used += sum(count(m.content) + 4 for m in history)    return used + reserve_for_reply <= model_window

The reserve_for_reply parameter is the piece people forget. The context window holds input and output. Filling 127,500 tokens of a 128,000-token window leaves 500 tokens for the answer, and you get a response cut off mid-sentence with no error — the same silent truncation that produces "the model got worse" bug reports.

Always reserve output space before you fill the context window. A request that fits and an answer that fits are two different constraints, and only one of them fails loudly.

Fitting more in than fits: three strategies

Sliding window — keep the most recent

Python
def window(messages, budget_tokens: int):    kept, used = [], 0    for m in reversed(messages):                 # newest first        c = count(m.content) + 4        if used + c > budget_tokens:            break        kept.append(m); used += c    kept.reverse()    while kept and kept[0].type == "ai":         # never start on an assistant turn        kept.pop(0)    return kept

Cheap, deterministic, no extra API call. The cost is total amnesia about anything that falls out — the customer's name mentioned in turn two vanishes at turn thirty. Right for chat where recency dominates, wrong when early facts matter.

Summarisation — compress the old

Python
SUMMARISE = ChatPromptTemplate.from_template(    "Compress the exchange below into at most 150 words. Preserve every name, "    "identifier, number, date and commitment exactly. Drop pleasantries and "    "restatements. Write in the third person.\n\n{chunk}")def compress(messages, keep_recent=8):    if len(messages) <= keep_recent + 4:        return messages    old, recent = messages[:-keep_recent], messages[-keep_recent:]    text = "\n".join(f"{m.type}: {m.content}" for m in old)    summary = (SUMMARISE | cheap_model | StrOutputParser()).invoke({"chunk": text})    return [SystemMessage(content=f"Summary of earlier conversation:\n{summary}")] + recent

The arithmetic that justifies it: 24 old messages averaging 120 tokens is 2,880 tokens, compressed to about 200 — a 93% reduction on that portion. The costs are an extra API call, a few hundred milliseconds, and irreversible information loss. The explicit instruction to preserve identifiers is what stops the loss being catastrophic; a summariser told merely to "summarise" reliably drops the order number, which is precisely the token you needed.

Retrieval — fetch only what is relevant

The strongest option when the source is a corpus rather than a conversation. Instead of choosing which parts of a 300-page document to keep, index it in chunks and retrieve the four chunks that match the question. Context stays roughly constant no matter how large the corpus grows.

StrategyExtra costLoses informationBest for
Sliding windowNoneYes, silentlyChat where only recency matters
SummarisationOne cheap call per compressionYes, controllablyLong conversations with facts worth keeping
RetrievalEmbedding + index infrastructureOnly what is not retrievedLarge static corpora, document Q&A
Structured state alongsideNegligibleNo, for tracked fieldsFacts your code depends on being exact

These combine, and in production they usually should: a structured state object for the handful of facts that must never be lost, a summary of the middle, the last few turns verbatim, and retrieval for the reference material.

A complete extraction, end to end

Pulling both halves together — bounded input, guaranteed output. The task: extract action items from a meeting transcript that may be far too long for one request.

Python
from typing import Literal, Optionalfrom pydantic import BaseModel, Fieldclass ActionItem(BaseModel):    task: str = Field(description="What must be done, as an imperative phrase")    owner: Optional[str] = Field(        default=None, description="Person responsible. Null if not assigned aloud.")    due: Optional[str] = Field(        default=None, description="Deadline as YYYY-MM-DD. Null if none was stated.")    priority: Literal["high", "medium", "low"] = Field(        description="Inferred urgency from the discussion")    quote: str = Field(description="The sentence from the transcript that states this item")class Extraction(BaseModel):    items: list[ActionItem] = Field(default_factory=list)    decisions: list[str] = Field(default_factory=list,                                 description="Decisions reached, one sentence each")extractor = model.with_structured_output(Extraction)def extract(transcript: str, window_tokens: int = 100_000) -> Extraction:    if count(transcript) <= window_tokens:        return extractor.invoke(PROMPT.format(text=transcript))    # Too long: process in overlapping chunks and merge    merged = Extraction()    for chunk in chunk_by_tokens(transcript, size=window_tokens, overlap=2_000):        part = extractor.invoke(PROMPT.format(text=chunk))        merged.items.extend(part.items)        merged.decisions.extend(part.decisions)    seen, unique = set(), []    for it in merged.items:        key = it.task.lower().strip()        if key not in seen:            seen.add(key); unique.append(it)    merged.items = unique    return merged

Three deliberate choices. The quote field forces the model to ground each item in actual transcript text, which makes hallucinated action items both rarer and instantly checkable — you can assert that item.quote appears in the source and reject the ones that do not. The 2,000-token overlap between chunks stops an action item that straddles a boundary from being lost by both sides. And the deduplication step exists because of the overlap: items in the shared region get extracted twice, and merging without dedup silently doubles them.

Where people get this wrong

Believing JSON mode enforces a schema. It enforces syntax. Failures 3 and 4 from the opening pass JSON mode cleanly and corrupt your data anyway.

Making every field required. The fastest way to manufacture hallucinations. If the document might not contain a due date, the schema must permit its absence — otherwise the model fabricates one, and it will be plausible enough that nobody notices for months.

Writing schemas without descriptions. A bare due_date: str gets you three date formats in one column. The description is the instruction; the field name alone is not.

Estimating tokens with len(text) / 4 and then enforcing a limit with it. Fine for a rough plan, wrong for a gate. Code, non-English text and long identifiers all tokenise far worse than the rule of thumb — code is often closer to 2.5 characters per token — so the estimate under-counts exactly where the input is largest.

Forgetting the reply budget. Filling the window to the brim gives you a truncated answer and no exception.

Using one provider's tokeniser to budget another provider's calls. Different vocabularies, different counts, wrong bill.

Chunking without overlap. Every fact that crosses a chunk boundary is lost, and it is lost silently — the extraction looks complete.

What this changes about how you build

Structured output moves the failure from runtime to design time, and that is the whole point. Before it, a model returning "PAID" instead of "paid" is a production incident discovered by a downstream report; after it, that value cannot be produced at all, because Literal["paid", "unpaid", "overdue", "unknown"] forbids it. You have converted a class of bug into a class of impossibility.

The practical consequence is that the schema becomes the specification of the feature. When someone asks for a new field, the conversation is about the type, whether it is optional, what its allowed values are and how to describe it — which is a far more productive conversation than "make the prompt better". Write the schema first, before the prompt, and the prompt often turns out to be two sentences.

The context side has a matching consequence: measure before you optimise. Log the input token count of every call from day one. Almost every team that hits a context limit discovers the cause is not the user's document but their own accumulated system prompt, a tool returning raw JSON, or a history that was never trimmed. Those are cheap fixes, but only if you have the number that points at them. Without it you will reach for a bigger model, pay several times more, and hit the same wall three months later with the same underlying cause.