Course Content
Prompt Engineering Mastery
6 sections · 32 lessons
How can you design a prompt to extract structured data from messy text?
What you need to know
Messy text — emails, WhatsApp messages, OCR output — has the facts you need in random order and format. The goal is a fixed shape that code can store.
The five parts of a reliable extraction prompt
- A schema — each field, its type and its format (
YYYY-MM-DD, amount as a number without commas or currency symbol). - A missing-value rule — "use null if the text does not state it; do not infer."
- Delimiters — the input inside
<email>tags, marked as data. - Normalisation rules — "Rs", "INR" and "₹" all become a number in rupees.
- Examples for hard cases — two amounts, a date written as "3rd Aug".
Structured outputs
Before this feature existed, teams asked politely for JSON, or prefilled the assistant's reply with {, then parsed with regex and retried. Prefilling is not supported on the newest Claude models, and structured outputs make most of that code unnecessary.
1from datetime import date2from pydantic import BaseModel3import anthropic45class Payment(BaseModel):6 payer_name: str | None7 phone: str | None8 amount_inr: float | None9 paid_on: date | None10 plan: str | None1112PROMPT = "Extract the payment described in <email>. Use null for anything not stated."13email = open("payment_email.txt").read()1415client = anthropic.Anthropic()16response = client.messages.parse(17 model="claude-opus-5",18 max_tokens=1024,19 messages=[{"role": "user", "content": PROMPT + "\n<email>\n" + email + "\n</email>"}],20 output_format=Payment,21)22payment = response.parsed_output # a validated Payment objectThe Pydantic model is the schema; the SDK sends it and returns a validated object. Structured outputs guarantee the shape, not the truth — a wrong amount in a valid field still passes, so business checks stay in your code.
A real-life example
A subscription business gets payment confirmations by email. The first prompt:
Pull out the important details from this email.Hi, this is Ravi (98XXXXXX21). Paid Rs 2,400 on 3rd Aug for the annualplan. Last year I paid 2,100.The output is a friendly paragraph, and when forced into JSON it sometimes picks 2,100. The improved prompt:
Extract the payment described in <email>. The payment is the one thecustomer says they just made, not past payments.Rules:- amount_inr: a number, no commas or currency symbols.- paid_on: YYYY-MM-DD; assume the year 2026 if missing.- Use null for anything not stated. Do not guess.<email>Hi, this is Ravi (98XXXXXX21). Paid Rs 2,400 on 3rd Aug for the annualplan. Last year I paid 2,100.</email>With the schema passed as a structured output, the reply is always {"payer_name": "Ravi", "phone": "98XXXXXX21", "amount_inr": 2400, "paid_on": "2026-08-03", "plan": "annual"}. Code then checks that 2,400 matches the annual plan's price; mismatches go to a person. Wrong amounts fall from 5% to under 1%.
Follow-up questions to expect
- "Why still validate if the API guarantees the schema?" — The schema guarantees types and fields, not correct values. Only your business rules catch a wrong but well-formed number.
- "What if the output is cut off?" — A reply that hits the token limit or is refused may not complete the schema. Check the stop reason before parsing.
- "How do you handle very long documents?" — Chunk them, extract per chunk with source spans, then merge and de-duplicate in code.